Practical Real-time Data Processing and Analytics

Book Image

Practical Real-time Data Processing and Analytics

Book Image

Practical Real-time Data Processing and Analytics

Overview of this book

With the rise of Big Data, there is an increasing need to process large amounts of data continuously, with a shorter turnaround time. Real-time data processing involves continuous input, processing and output of data, with the condition that the time required for processing is as short as possible. This book covers the majority of the existing and evolving open source technology stack for real-time processing and analytics. You will get to know about all the real-time solution aspects, from the source to the presentation to persistence. Through this practical book, you’ll be equipped with a clear understanding of how to solve challenges on your own. We’ll cover topics such as how to set up components, basic executions, integrations, advanced use cases, alerts, and monitoring. You’ll be exposed to the popular tools used in real-time processing today such as Apache Spark, Apache Flink, and Storm. Finally, you will put your knowledge to practical use by implementing all of the techniques in the form of a practical, real-world use case. By the end of this book, you will have a solid understanding of all the aspects of real-time data processing and analytics, and will know how to deploy the solutions in production environments in the best possible manner.

Title Page

Credits

About the Authors

About the Authors

About the Reviewers

About the Reviewers

www.PacktPub.com

www.PacktPub.com

Customer Feedback

Customer Feedback

Preface

Free Chapter

Introducing Real-Time Analytics

Introducing Real-Time Analytics

What is big data?

Big data infrastructure

Real–time analytics – the myth and the reality

Near real–time solution – an architecture that works

Lambda architecture – analytics possibilities

IOT – thoughts and possibilities

Cloud – considerations for NRT and IOT

Real Time Applications – The Basic Ingredients

Real Time Applications – The Basic Ingredients

The NRT system and its building blocks

NRT – high-level system view

NRT – technology view

Understanding and Tailing Data Streams

Understanding and Tailing Data Streams

Understanding data streams

Setting up infrastructure for data ingestion

Taping data from source to the processor - expectations and caveats

Comparing and choosing what works best for your use case

Setting up the Infrastructure for Storm

Setting up the Infrastructure for Storm

Overview of Storm

Storm architecture and its components

Setting up and configuring Storm

Real-time processing job on Storm

Configuring Apache Spark and Flink

Configuring Apache Spark and Flink

Setting up and a quick execution of Spark

Setting up and a quick execution of Flink

Setting up and a quick execution of Apache Beam

Balancing in Apache Beam

Integrating Storm with a Data Source

Integrating Storm with a Data Source

RabbitMQ – messaging that works

RabbitMQ exchanges

RabbitMQ – integration with Storm

PubNub data stream publisher

String together Storm-RMQ-PubNub sensor data topology

From Storm to Sink

From Storm to Sink

Setting up and configuring Cassandra

Storm and Cassandra topology

Storm and IMDB integration for dimensional data

Integrating the presentation layer with Storm

Storm Trident

State retention and the need for Trident

Basic Storm Trident topology

Trident internals

Trident operations

Working with Spark

Working with Spark

Distinct advantages of Spark

Spark – use cases

Spark architecture - working inside the engine

Spark pragmatic concepts

Spark 2.x – advent of data frames and datasets

Working with Spark Operations

Working with Spark Operations

Spark – packaging and API

RDD pragmatic exploration

Shared variables – broadcast variables and accumulators

Spark Streaming

Spark Streaming

Spark Streaming concepts

Spark Streaming - introduction and architecture

Packaging structure of Spark Streaming

Connecting Kafka to Spark Streaming

Working with Apache Flink

Working with Apache Flink

Flink architecture and execution engine

Flink basic components and processes

Integration of source stream to Flink

Flink processing and computation

Flink persistence

Case Study

Tools and frameworks

Setting up the infrastructure

Implementing the case study

Running the case study

Customer Reviews

5 star

0

4 star

0

3 star

0

2 star

0

1 star

0

Spark 2.x – advent of data frames and datasets

With Spark 2.x we have two new spark computational abstractions:

Data frames: These are distributed, resilient, fault tolerant in-memory data structures that are capable of handling only structured data, which means they are designed to manage data that can be segregated in fixed typed columns. Though it may sound like a limitation with respect to RDD, which can handle any type of unstructured data, in practical terms this structured abstraction over the data makes it very easy to manipulate and work over a large volume of structured data, the way we used to with RDBMS.
Datasets: It's an extension of the Spark data frame. It's a type safe object-oriented interface. For the sake of simplicity, one could say that data frames are actually an un-typed dataset. This newest API in spark pragmatic abstraction actually leverages features of tungsten in-memory encoding and catalysts optimizer.