Building Big Data Pipelines with Apache Beam

By : Jan Lukavský

Building Big Data Pipelines with Apache Beam

By: Jan Lukavský

Overview of this book

Apache Beam is an open source unified programming model for implementing and executing data processing pipelines, including Extract, Transform, and Load (ETL), batch, and stream processing. This book will help you to confidently build data processing pipelines with Apache Beam. You’ll start with an overview of Apache Beam and understand how to use it to implement basic pipelines. You’ll also learn how to test and run the pipelines efficiently. As you progress, you’ll explore how to structure your code for reusability and also use various Domain Specific Languages (DSLs). Later chapters will show you how to use schemas and query your data using (streaming) SQL. Finally, you’ll understand advanced Apache Beam concepts, such as implementing your own I/O connectors. By the end of this book, you’ll have gained a deep understanding of the Apache Beam model and be able to apply it to solve problems.

Preface

Who this book is for

What this book covers

To get the most out of this book

Download the example code files

Download the color images

Conventions used

Get in touch

Share Your Thoughts

Section 1 Apache Beam: Essentials

Free Chapter

Chapter 1: Introduction to Data Processing with Apache Beam

Technical requirements

Why Apache Beam?

Writing your first pipeline

Running our pipeline against streaming data

Exploring the key properties of unbounded data

Measuring event time progress inside data streams

Assigning data to windows

Unifying batch and streaming data processing

Summary

Chapter 2: Implementing, Testing, and Deploying Basic Pipelines

Technical requirements

Setting up the environment for this book

Task 1 – Calculating the K most frequent words in a stream of lines of text

Task 2 – Calculating the maximal length of a word in a stream

Specifying the PCollection Coder object and the TypeDescriptor object

Understanding default triggers, on time, and closing behavior

Introducing the primitive PTransform object – Combine

Task 3 – Calculating the average length of words in a stream

Task 4 – Calculating the average length of words in a stream with fixed lookback

Ensuring pipeline upgradability

Task 5 – Calculating performance statistics for a sport activity tracking application

Introducing the primitive PTransform object – GroupByKey

Introducing the primitive PTransform object – Partition

Summary

Chapter 3: Implementing Pipelines Using Stateful Processing

Technical requirements

Task 6 – Using an external service for data augmentation

Introducing the primitive PTransform object – stateless ParDo

Task 7 – Batching queries to an external RPC service

Task 8 – Batching queries to an external RPC service with defined batch sizes

Introducing the primitive PTransform object – stateful ParDo

Using side outputs

Defining droppable data in Beam

Task 9 – Separating droppable data from the rest of the data processing

Task 10 – Separating droppable data from the rest of the data processing, part 2

Using side inputs

Summary

Section 2 Apache Beam: Toward Improving Usability

Chapter 4: Structuring Code for Reusability

Technical requirements

Explaining PTransform expansion

Task 11 – Enhancing SportTracker by runner motivation using side inputs

Introducing composite transform – CoGroupByKey

Task 12 – enhancing SportTracker by runner motivation using CoGroupByKey

Introducing the Join library DSL

Stream-to-stream joins explained

Task 13 – Writing a reusable PTransform – StreamingInnerJoin

Table-stream duality

Summary

Chapter 5: Using SQL for Pipeline Implementation

Technical requirements

Understanding schemas

Implementing our first streaming pipeline using SQL

Task 14 – Implementing SQLMaxWordLength

Task 15 – Implementing SchemaSportTracker

Task 16 – Implementing SQLSportTrackerMotivation

Further development of Apache Beam SQL

Summary

Chapter 6: Using Your Preferred Language with Portability

Technical requirements

Introducing the portability layer

Implementing our first pipelines in the Python SDK

Task 17 – Implementing MaxWordLength in the Python SDK

Python SDK type hints and coders

Task 18 – Implementing SportTracker in the Python SDK

Task 19 – Implementing RPCParDo in the Python SDK

Task 20 – Implementing SportTrackerMotivation in the Python SDK

Using the DataFrame API

Interactive programming using InteractiveRunner

Introducing and using cross-language pipelines

Summary

Section 3 Apache Beam: Advanced Concepts

Chapter 7: Extending Apache Beam's I/O Connectors

Technical requirements

Defining splittable DoFn as a unification for bounded and unbounded sources

Task 21 – Implementing our own splittable DoFn – a streaming file source

Task 22 – A non-I/O application of splittable DoFn – PiSampler

The legacy Source API and the Read transform

Writing a custom data sink

Summary

Chapter 8: Understanding How Runners Execute Pipelines

Describing the anatomy of an Apache Beam runner

Explaining the differences between classic and portable runners

Understanding how a runner handles state

Exploring the Apache Beam capability matrix

Understanding windowing semantics in depth

Debugging pipelines and using Apache Beam metrics for observability

Summary

Why subscribe?

Other Books You May Enjoy

Packt is searching for authors like you

Share Your Thoughts

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Assigning data to windows

We have already touched on (but have not yet defined) the concept of a window. A window is a specific, bounded range of data within a data stream. Beam has several types of pre-defined window functions:

Tumbling windows
Sliding windows
Session windows

Tumbling windows are for assigning data elements into a single window of a pre-defined length, as follows:

Figure 1.9 – Tumbling windows

Tumbling windows can each have exactly the same fixed length (for example, 1 hour or 1 day in what are called fixed windows), or different lengths (for example, 1 month in what are called calendar windows). The common property of tumbling windows is that the event time of each element can be assigned to exactly one window, and that these windows cover a continuous, (possibly) infinite time range, without any gaps.

Sliding windows are windows that assign data elements into multiple windows, shifted by a time period called a slide, as shown in the following figure:

Figure 1.10 – Sliding windows

Sliding windows have the same fixed window length (for example, 1 hour) and the same fixed slide (for example, 10 minutes). A sliding window of 1 hour with a slide of 10 minutes assigns each event time into six distinct windows, each shifted by 10 minutes.

The last type of window is called a session window. This type of window is special in several ways. Unlike both previous types, session windows are key unaligned. What does that mean? Neither tumbling nor sliding windows depend on the data itself – each data element is assigned to a window (or several windows) based solely on the element's timestamp. The boundary of all the windows in the stream is exactly aligned for all the data. This is not the case for session windows. Session windows split the stream into independent sub-streams based on a user-provided key for each element in the stream. We can imagine the key as a color representing each stream element. Session windows group only elements having the same color, therefore, windows in the stream are no longer aligned on the same boundary. We can illustrate this as follows:

Figure 1.11 – Session windows

We can see from Figure 1.11 that different keys (types) of elements are grouped in different windows. There, one other parameter that has to be specified: a session gap duration. This duration is a timeout (in the event time) that has to elapse between the timestamps of two successive elements with the same key in order to prevent assigning them in the same window. That is to say, as long as elements for a key arrive with a frequency higher than the gap duration, all are placed in the same window. Once there is a delay of at least the gap duration, the window is closed, and another window will be created when a new element arrives. This type of window is frequently used when analyzing user sessions in web clickstreams (which is where the name session window came from).

There is one more special window type called a global window. This very special type of window assigns all data elements into a single window, regardless of their timestamp. Therefore, the window spans a complete time interval from –infinity to +infinity. This window is used as a default window before any other window is applied. We'll look into this later in this chapter.

Defining the life cycle of a state in terms of windows

Windows are actually a way of scoping a state in computation. Each state is valid within the context of a window, and each window has its own independent state.

Figure 1.12 illustrates state scoping:

Figure 1.12 – Scoping state within windows

The scoping of states by windows brings up another crucial concept of stream processing: late data elements. One such element is shown in Figure 1.7.

We can state the problem as follows: when can we clear and discard the state that belongs to a particular window? Obviously, it is impractical to keep all states of all windows open forever, because each window carries a non-zero memory footprint, and keeping the window around for an unbounded time would cause the memory to be depleted over time. On the other hand, deleting the state right after the watermark passes the timestamp that marks the end of the window would mean we need a perfect watermark (a watermark that never produces late data). Any possible late data would mean we would produce incorrect outputs – the state would be cleared before all the data elements belonging to the respective window could be processed and therefore would have to be dropped or would produce a completely wrong outcome.

One option would be to define semantics that would require the watermark to advance only when the probability of late data is sufficiently low. We would drop all data that arrived after the watermark and pretend that we didn't see it. If the watermark produces a sufficiently low number of this late data, the error introduced by dropping the late data would be negligible. The crucial problem with this approach is that it necessarily introduces very high latency due to the out-of-orderness of stream processing. We would therefore face a latency versus correctness trade-off, when our goal ideally should be to have both high correctness and low latency.

To resolve this dilemma, stream processing engines introduce an additional concept called allowed lateness. This defines a timeout (in the event time) after which the state in a window can be cleared and all remaining data can be cleared. This option gives us the possibility to achieve the following:

Enable the watermark heuristic to advance sufficiently quickly to not incur unnecessary latency.
Enable an independent measure of how many states are to be kept around, even after their maximal timestamp has already passed.

We illustrate this concept in Figure 1.13, which shows a simple watermark heuristic that just shifts the processing time by a constant duration (which will define minimal latency) and a late data boundary, which shifts the watermark by an additional allowed lateness duration. This might introduce data that will be actually dropped but can now be tuned independently:

Figure 1.13 – Allowed lateness

Important note

Practical watermark implementations do not typically use a fixed shift between the watermark and processing time, but rather use statistics inferred from consumed data to produce a watermark that is non-linear in terms of the processing time.

The definition of on-time and late data brings up one last technical term that appears in the context of triggers (see the States, triggers, and timers section as a reminder). When a trigger condition is met and the trigger causes output data to be emitted downstream, three possible conditions can occur:

The watermark has not yet reached the end timestamp of a window.
The watermark has crossed the window end timestamp, and this is the first activation of a trigger since then.
The watermark has passed the window end timestamp, and this is not the first activation of a trigger.

According to these three conditions, we can mark the resulting downstream data element as one of the following:

Early: The data is emitted prior to terminating the respective window's end timestamp – this means that we output speculative partial results.
On-time: This marks data that was calculated once the window's end timestamp was reached.
Late: This contains any output with late data incorporated.

Beam calls data emitted as a result of trigger firings a pane and puts the information about lateness or earliness of such firing into the PaneInfo object.

Pane accumulation

When a trigger fires and causes data to be output from the current window(s) to downstream processing, there are several options that can be used with both the state associated with the window and with the resulting value itself.

After a trigger fires and data is output downstream, we have essentially two options:

Reset the state to an empty (initial) state (discard).
Keep the state intact (accumulate).

This concept might be a little confusing, so we'll demonstrate it with an example. Let's assume that we want to count the number of elements in a stream every minute in the processing time. In general, window functions are based on event time, so to get something that would resemble a processing time window, we can use the following:

// Window into single window and specify trigger
PCollection<String> windowed =
    words.apply(
        Window.<String>into(new GlobalWindows())
          .triggering(
           Repeatedly.forever(
              AfterProcessingTime.pastFirstElementInPane()
               .plusDelayOf(Duration.standardSeconds(1))))
          .discardingFiredPanes());

Please investigate the complete source code in the com.packtpub.beam.chapter1.ProcessingTimeWindow class.

We can run this pipeline using the following:

chapter1$ ../mvnw exec:java \
 -Dexec.mainClass=com.packtpub.beam.chapter1.ProcessingTimeWindow

Please feel free to experiment with changing discardingFiredPanes to accumulatingFiredPanes to see how the output differs. In the accumulation mode, the output contains the sum of the elements from the beginning, while in the discarding mode, it contains only increments from the last trigger firing.

Now that we have discussed all the key properties of data streams, let's see how we can use this knowledge to close the gap between batch processing and real-time stream processing!

Building Big Data Pipelines with Apache Beam

By : Jan Lukavský

Building Big Data Pipelines with Apache Beam

By: Jan Lukavský

Overview of this book

Related Content you might be interested in

Current Title:

Building Big Data Pipelines with Apache Beam

Learning Apache Apex

Practical Real-time Data Processing and Analytics

Assigning data to windows

Defining the life cycle of a state in terms of windows

Pane accumulation