Practical Real-time Data Processing and Analytics

Practical Real-time Data Processing and Analytics

Overview of this book

With the rise of Big Data, there is an increasing need to process large amounts of data continuously, with a shorter turnaround time. Real-time data processing involves continuous input, processing and output of data, with the condition that the time required for processing is as short as possible. This book covers the majority of the existing and evolving open source technology stack for real-time processing and analytics. You will get to know about all the real-time solution aspects, from the source to the presentation to persistence. Through this practical book, you’ll be equipped with a clear understanding of how to solve challenges on your own. We’ll cover topics such as how to set up components, basic executions, integrations, advanced use cases, alerts, and monitoring. You’ll be exposed to the popular tools used in real-time processing today such as Apache Spark, Apache Flink, and Storm. Finally, you will put your knowledge to practical use by implementing all of the techniques in the form of a practical, real-world use case. By the end of this book, you will have a solid understanding of all the aspects of real-time data processing and analytics, and will know how to deploy the solutions in production environments in the best possible manner.

Title Page

Credits

About the Authors

About the Reviewers

www.PacktPub.com

Customer Feedback

Preface

Free Chapter

Introducing Real-Time Analytics

What is big data?

Big data infrastructure

Real–time analytics – the myth and the reality

Near real–time solution – an architecture that works

Lambda architecture – analytics possibilities

IOT – thoughts and possibilities

Cloud – considerations for NRT and IOT

Summary

Real Time Applications – The Basic Ingredients

The NRT system and its building blocks

NRT – high-level system view

NRT – technology view

Summary

Understanding and Tailing Data Streams

Understanding data streams

Setting up infrastructure for data ingestion

Taping data from source to the processor - expectations and caveats

Comparing and choosing what works best for your use case

Do it yourself

Summary

Setting up the Infrastructure for Storm

Overview of Storm

Storm architecture and its components

Setting up and configuring Storm

Real-time processing job on Storm

Summary

Configuring Apache Spark and Flink

Setting up and a quick execution of Spark

Setting up and a quick execution of Flink

Setting up and a quick execution of Apache Beam

Balancing in Apache Beam

Summary

Integrating Storm with a Data Source

RabbitMQ – messaging that works

RabbitMQ exchanges

RabbitMQ – integration with Storm

PubNub data stream publisher

String together Storm-RMQ-PubNub sensor data topology

Summary

From Storm to Sink

Setting up and configuring Cassandra

Storm and Cassandra topology

Storm and IMDB integration for dimensional data

Integrating the presentation layer with Storm

Do It Yourself

Summary

Storm Trident

State retention and the need for Trident

Basic Storm Trident topology

Trident internals

Trident operations

DRPC

Do It Yourself

Summary

Working with Spark

Spark overview

Distinct advantages of Spark

Spark – use cases

Spark architecture - working inside the engine

Spark pragmatic concepts

Spark 2.x – advent of data frames and datasets

Summary

Working with Spark Operations

Spark – packaging and API

RDD pragmatic exploration

Shared variables – broadcast variables and accumulators

Summary

Spark Streaming

Spark Streaming concepts

Spark Streaming - introduction and architecture

Packaging structure of Spark Streaming

Connecting Kafka to Spark Streaming

Summary

Working with Apache Flink

Flink architecture and execution engine

Flink basic components and processes

Integration of source stream to Flink

Flink processing and computation

Flink persistence

FlinkCEP

Pattern API

Gelly

DIY

Summary

Case Study

Introduction

Data modeling

Tools and frameworks

Setting up the infrastructure

Implementing the case study

Running the case study

Summary

Customer Reviews

5 star

4 star

3 star

2 star

1 star

What is big data?

Well to begin with, in simple terms, big data helps us deal with three V's – volume, velocity, and variety. Recently, two more V's were added to it, making it a five–dimensional paradigm; they are veracity and value

Volume: This dimension refers to the amount of data; look around you, huge amounts of data are being generated every second – it may be the email you send, Twitter, Facebook, or other social media, or it can just be all the videos, pictures, SMS messages, call records, and data from varied devices and sensors. We have scaled up the data–measuring metrics to terabytes, zettabytes and Yottabytes – they are all humongous figures. Look at Facebook alone; it's like ~10 billion messages on a day, consolidated across all users. We have ~5 billion likes a day and around ~400 million photographs are uploaded each day. Data statistics in terms of volume are startling; all of the data generated from the beginning of time to 2008 is kind of equivalent to what we generate in a day today, and I am sure soon it will be an hour. This volume aspect alone is making the traditional database dwarf to store and process this amount of data in reasonable and useful time frames, though a big data stack can be employed to store process and compute on amazingly large data sets in a cost–effective, distributed, and reliably efficient manner.
Velocity: This refers to the data generation speed, or the rate at which data is being generated. In today's world, where we mentioned that the volume of data has undergone a tremendous surge, this aspect is not lagging behind. We have loads of data because we are able to generate it so fast. Look at social media; things are circulated in seconds and they become viral, and the insight from social media is analysed in milliseconds by stock traders, and that can trigger lots of activity in terms of buying or selling. At a target point of sale counter it takes a few seconds for a credit card swipe, and within that fraudulent transaction processing, payment, bookkeeping, and acknowledgement is all done. Big data gives us the power to analyse the data at tremendous speed.
Variety: This dimension tackles the fact that the data can be unstructured. In the traditional database world, and even before that, we were used to having a very structured form of data that fitted neatly into tables. Today, more than 80% of data is unstructured – quotable examples are photos, video clips, social media updates, data from variety of sensors, voice recordings, and chat conversations. Big data lets you store and process this unstructured data in a very structured manner; in fact, it effaces the variety.
Veracity: It's all about validity and correctness of data. How accurate and usable is the data? Not everything out of millions and zillions of data records is corrected, accurate, and referable. That's what actual veracity is: how trustworthy the data is and what the quality of the data is. Examples of data with veracity include Facebook and Twitter posts with nonstandard acronyms or typos. Big data has brought the ability to run analytics on this kind of data to the table. One of the strong reasons for the volume of data is veracity.
Value: This is what the name suggests: the value that the data actually holds. It is unarguably the most important V or dimension of big data. The only motivation for going towards big data for processing super large data sets is to derive some valuable insight from it. In the end, it's all about cost and benefits.

Big data is a much talked about technology across businesses and the technical world today. There are myriad domains and industries that are convinced of its usefulness, but the implementation focus is primarily application-oriented, rather than infrastructure-oriented. The next section predominantly walks you through the same.

Practical Real-time Data Processing and Analytics

Practical Real-time Data Processing and Analytics

Overview of this book

Related Content you might be interested in

Current Title:

Practical Real-time Data Processing and Analytics

Mastering Apache Storm

Building Data Streaming Applications with Apache Kafka

Apache Kafka 1.0 Cookbook

What is big data?