Practical Real-time Data Processing and Analytics

Practical Real-time Data Processing and Analytics

Overview of this book

With the rise of Big Data, there is an increasing need to process large amounts of data continuously, with a shorter turnaround time. Real-time data processing involves continuous input, processing and output of data, with the condition that the time required for processing is as short as possible. This book covers the majority of the existing and evolving open source technology stack for real-time processing and analytics. You will get to know about all the real-time solution aspects, from the source to the presentation to persistence. Through this practical book, you’ll be equipped with a clear understanding of how to solve challenges on your own. We’ll cover topics such as how to set up components, basic executions, integrations, advanced use cases, alerts, and monitoring. You’ll be exposed to the popular tools used in real-time processing today such as Apache Spark, Apache Flink, and Storm. Finally, you will put your knowledge to practical use by implementing all of the techniques in the form of a practical, real-world use case. By the end of this book, you will have a solid understanding of all the aspects of real-time data processing and analytics, and will know how to deploy the solutions in production environments in the best possible manner.

Title Page

Credits

About the Authors

About the Reviewers

www.PacktPub.com

Customer Feedback

Preface

Free Chapter

Introducing Real-Time Analytics

What is big data?

Big data infrastructure

Real–time analytics – the myth and the reality

Near real–time solution – an architecture that works

Lambda architecture – analytics possibilities

IOT – thoughts and possibilities

Cloud – considerations for NRT and IOT

Summary

Real Time Applications – The Basic Ingredients

The NRT system and its building blocks

NRT – high-level system view

NRT – technology view

Summary

Understanding and Tailing Data Streams

Understanding data streams

Setting up infrastructure for data ingestion

Taping data from source to the processor - expectations and caveats

Comparing and choosing what works best for your use case

Do it yourself

Summary

Setting up the Infrastructure for Storm

Overview of Storm

Storm architecture and its components

Setting up and configuring Storm

Real-time processing job on Storm

Summary

Configuring Apache Spark and Flink

Setting up and a quick execution of Spark

Setting up and a quick execution of Flink

Setting up and a quick execution of Apache Beam

Balancing in Apache Beam

Summary

Integrating Storm with a Data Source

RabbitMQ – messaging that works

RabbitMQ exchanges

RabbitMQ – integration with Storm

PubNub data stream publisher

String together Storm-RMQ-PubNub sensor data topology

Summary

From Storm to Sink

Setting up and configuring Cassandra

Storm and Cassandra topology

Storm and IMDB integration for dimensional data

Integrating the presentation layer with Storm

Do It Yourself

Summary

Storm Trident

State retention and the need for Trident

Basic Storm Trident topology

Trident internals

Trident operations

DRPC

Do It Yourself

Summary

Working with Spark

Spark overview

Distinct advantages of Spark

Spark – use cases

Spark architecture - working inside the engine

Spark pragmatic concepts

Spark 2.x – advent of data frames and datasets

Summary

Working with Spark Operations

Spark – packaging and API

RDD pragmatic exploration

Shared variables – broadcast variables and accumulators

Summary

Spark Streaming

Spark Streaming concepts

Spark Streaming - introduction and architecture

Packaging structure of Spark Streaming

Connecting Kafka to Spark Streaming

Summary

Working with Apache Flink

Flink architecture and execution engine

Flink basic components and processes

Integration of source stream to Flink

Flink processing and computation

Flink persistence

FlinkCEP

Pattern API

Gelly

DIY

Summary

Case Study

Introduction

Data modeling

Tools and frameworks

Setting up the infrastructure

Implementing the case study

Running the case study

Summary

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Cloud – considerations for NRT and IOT

Cloud is nothing but a term used to identify a capability where computational capability is available over internet. We have all been acquainted with physical machines, servers, and data centres. The advent of the cloud has taken us to a world of virtualization where we are moving out to virtual nodes, virtualized clusters, and even virtual data centers. Now, I can have virtualization of hardware to play with and have my clusters built using VMs spawned over a few actual machines. So, it's like having software at play over physical hardware. The next step was the cloud, where we have all virtual compute capability hosted and available over the net.

The services that are part of the cloud bouquet are:

Infrastructure as a Service (IaaS): It's basically a cloud variant of fundamental physical computers. It actually replaces the actual machines, the servers and hardware storage and the networking by a virtualization layer operating over the net. The IaaS lets you build this entire virtual infrastructure, where it's actually software that's imitating actual hardware.
Platform as a Service (PaaS): Once we have sorted out the hardware virtualization part, the next obvious step is to think about the next layer that operates over the raw computer hardware. This is the one that gels the programs with components like databases, servers, file storage, and so on. Here, for instance, if a database is exposed as PaaS then the programmer can use that as a service without worrying about the lower details like storage capacity, data protection, encryption, replication, and so on. Renowned examples of PaaS are Google App Engine and Heroku.
Software as a Service (SaaS): This one is the topmost layer in the stack of cloud computation; it's actually the layer that provides the solution as a service. These services are charged on a per user or per month basis and this model ensures that the end users have flexibility to enroll for and use the service without any license fee or locking period. Some of the most widely known, and typical, examples are Salesforce.com, and Google Apps.

Now that we have been introduced to and acquainted with cloud, the next obvious point to understand is what this buzz is all about and why is it that the advent of the cloud is closing curtains on the era of traditional data centers. Let's understand some of the key benefits of cloud computing that have actually made this platform a hot selling cake for NRT and IOT applications

It's on demand: The users can provision the computing components/resources as per the need and load. There is no need to make huge investments for the next X years on infrastructure in the name of future scaling and headroom. One can provision a cluster that's adequate to meet current requirements and that can then be scaled when needed by requesting more on–demand instances. So, the guarantee I get here as a user, is that I will get an instance when I demand for the same.
It lets us build truly elastic applications, which means that depending upon the load and the need, my deployments can scale up and scale down. This is a huge advantage and the way the cloud does it is very cost effective too. If I have an application which sees an occasional surge in traffic on the first of every month, then, on the cloud I don't have to provision the hardware required to meet the demand of surge on the first of the month for the entire 30 days. Instead, I can provision what I need on an average day and build a mechanism to scale up my cluster to meet the surge on the first and then automatically scale back to average size on the second of the month.
It's pay as you go: well, this is the most interesting feature of the cloud that beats the traditional hardware provisioning system, where to set up a data centers one has to plan up front for a huge investment. Cloud data centre, don't invite any such cost and one can pay only for the instances that are running, and that too, is generally handled on an hourly basis.

Practical Real-time Data Processing and Analytics

Practical Real-time Data Processing and Analytics

Overview of this book

Related Content you might be interested in

Current Title:

Practical Real-time Data Processing and Analytics

Mastering Apache Storm

Building Data Streaming Applications with Apache Kafka

Apache Kafka 1.0 Cookbook

Cloud – considerations for NRT and IOT