Hands-On Big Data Analytics with PySpark

By : Rudy Lai, Bartłomiej Potaczek

Hands-On Big Data Analytics with PySpark

By: Rudy Lai, Bartłomiej Potaczek

Overview of this book

Apache Spark is an open source parallel-processing framework that has been around for quite some time now. One of the many uses of Apache Spark is for data analytics applications across clustered computers. In this book, you will not only learn how to use Spark and the Python API to create high-performance analytics with big data, but also discover techniques for testing, immunizing, and parallelizing Spark jobs. You will learn how to source data from all popular data hosting platforms, including HDFS, Hive, JSON, and S3, and deal with large datasets with PySpark to gain practical big data experience. This book will help you work on prototypes on local machines and subsequently go on to handle messy data in production and at scale. This book covers installing and setting up PySpark, RDD operations, big data cleaning and wrangling, and aggregating and summarizing data into useful reports. You will also learn how to implement some practical and proven techniques to improve certain aspects of programming and administration in Apache Spark. By the end of the book, you will be able to build big data analytical solutions using the various PySpark offerings and also optimize them effectively.

Preface

Who this book is for

What this book covers

To get the most out of this book

Get in touch

Free Chapter

Installing Pyspark and Setting up Your Development Environment

An overview of PySpark

Setting up Spark on Windows and PySpark

Core concepts in Spark and PySpark

Summary

Getting Your Big Data into the Spark Environment Using RDDs

Loading data on to Spark RDDs

Parallelization with Spark RDDs

Basics of RDD operation

Summary

Big Data Cleaning and Wrangling with Spark Notebooks

Using Spark Notebooks for quick iteration of ideas

Sampling/filtering RDDs to pick out relevant data points

Splitting datasets and creating some new combinations

Summary

Aggregating and Summarizing Data into Useful Reports

Calculating averages with map and reduce

Faster average computations with aggregate

Pivot tabling with key-value paired data points

Summary

Powerful Exploratory Data Analysis with MLlib

Computing summary statistics with MLlib

Using Pearson and Spearman correlations to discover correlations

Testing our hypotheses on large datasets

Summary

Putting Structure on Your Big Data with SparkSQL

Manipulating DataFrames with Spark SQL schemas

Using Spark DSL to build queries

Summary

Transformations and Actions

Using Spark transformations to defer computations to a later time

Avoiding transformations

Using the reduce and reduceByKey methods to calculate the results

Performing actions that trigger computations

Reusing the same rdd for different actions

Summary

Immutable Design

Delving into the Spark RDD's parent/child chain

Using RDD in an immutable way

Using DataFrame operations to transform

Immutability in the highly concurrent environment

Using the Dataset API in an immutable way

Summary

Avoiding Shuffle and Reducing Operational Expenses

Detecting a shuffle in a process

Testing operations that cause a shuffle in Apache Spark

Changing the design of jobs with wide dependencies

Using keyBy() operations to reduce shuffle

Using a custom partitioner to reduce shuffle

Summary

Saving Data in the Correct Format

Saving data in plain text format

Leveraging JSON as a data format

Tabular formats – CSV

Using Avro with Spark

Columnar formats – Parquet

Summary

Working with the Spark Key/Value API

Available actions on key/value pairs

Using aggregateByKey instead of groupBy()

Actions on key/value pairs

Available partitioners on key/value data

Implementing a custom partitioner

Summary

Testing Apache Spark Jobs

Separating logic from Spark engine-unit testing

Integration testing using SparkSession

Mocking data sources using partial functions

Using ScalaCheck for property-based testing

Testing in different versions of Spark

Summary

Leveraging the Spark GraphX API

Creating a graph from a data source

Using the Vertex API

Using the Edge API

Calculating the degree of the vertex

Calculating PageRank

Summary

Other Books You May Enjoy

Leave a review - let other readers know what you think

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Core concepts in Spark and PySpark

Let's now look at the following core concepts in Spark and PySpark:

SparkContext
SparkConf
Spark shell

SparkContext

SparkContext is an object or concept within Spark. It is a big data analytical engine that allows you to programmatically harness the power of Spark.

The power of Spark can be seen when you have a large amount of data that doesn't fit into your local machine or your laptop, so you need two or more computers to process it. You also need to maintain the speed of processing this data while working on it. We not only want the data to be split among a few computers for computation; we also want the computation to be parallel. Lastly, you want this computation to look like one single computation.

Let's consider an example where we have a large contact database that has 50 million names, and we might want to extract the first name from each of these contacts. Obviously, it is difficult to fit 50 million names into your local memory, especially if each name is embedded within a larger contacts object. This is where Spark comes into the picture. Spark allows you to give it a big data file, and will help in handling and uploading this data file, while handling all the operations carried out on this data for you. This power is managed by Spark's cluster manager, as shown in the following diagram:

The cluster manager manages multiple workers; there could be 2, 3, or even 100. The main point is that Spark's technology helps in managing this cluster of workers, and you need a way to control how the cluster is behaving, and also pass data back and forth from the clustered rate.

A SparkContext lets you use the power of Spark's cluster manager as with Python objects. So with a SparkContext, you can pass jobs and resources, schedule tasks, and complete tasks the downstream from the SparkContext down to the Spark Cluster Manager, which will then take the results back from the Spark Cluster Manager once it has completed its computation.

Let's see what this looks like in practice and see how to set up a SparkContext:

First, we need to import SparkContext.
Create a new object in the sc variable standing for the SparkContext using the SparkContext constructor.
In the SparkContext constructor, pass a local context. We are looking at hands on PySpark in this context, as follows:

from pyspark import SparkContext
sc = SparkContext('local', 'hands on PySpark')

After we've established this, all we need to do is then use sc as an entry point to our Spark operation, as demonstrated in the following code snippet:

visitors = [10, 3, 35, 25, 41, 9, 29]
df_visitors = sc.parallelize(visitors)
df_visitors_yearly = df_visitors.map(lambda x: x*365).collect()
print(df_visitors_yearly)

Let's take an example; if we were to analyze the synthetic datasets of visitor counts to our clothing store, we might have a list of visitors denoting the daily visitors to our store. We can then create a parallelized version of the DataFrame, call sc.parallelize(visitors), and feed in the visitors datasets. df_visitors then creates for us a DataFrame of visitors. We can then map a function; for example, making the daily numbers and extrapolating them into a yearly number by mapping a lambda function that multiplies the daily number (x) by 365, which is the number of days in a year. Then, we call a collect() function to make sure that Spark executes on this lambda call. Lastly, we print out df_ visitors_yearly. Now, we have Spark working on this computation on our synthetic data behind the scenes, while this is simply a Python operation.

Spark shell

We will go back into our Spark folder, which is spark-2.3.2-bin-hadoop2.7, and start our PySpark binary by typing .\bin\pyspark.

We can see that we've started a shell session with Spark in the following screenshot:

Spark is now available to us as a spark variable. Let's try a simple thing in Spark. The first thing to do is to load a random file. In each Spark installation, there is a README.md markdown file, so let's load it into our memory as follows:

text_file = spark.read.text("README.md")

If we use spark.read.text and then put in README.md, we get a few warnings, but we shouldn't be too concerned about that at the moment, as we will see later how we are going to fix these things. The main thing here is that we can use Python syntax to access Spark.

What we have done here is put README.md as text data read by spark into Spark, and we can use text_file.count() can get Spark to count how many characters are in our text file as follows:

text_file.count()

From this, we get the following output:

We can also see what the first line is with the following:

text_file.first()

We will get the following output:

Row(value='# Apache Spark')

We can now count a number of lines that contain the word Spark by doing the following:

lines_with_spark = text_file.filter(text_file.value.contains("Spark"))

Here, we have filtered for lines using the filter() function, and within the filter() function, we have specified that text_file_value.contains includes the word "Spark", and we have put those results into the lines_with_spark variable.

We can modify the preceding command and simply add .count(), as follows:

text_file.filter(text_file.value.contains("Spark")).count()

We will now get the following output:

We can see that 20 lines in the text file contain the word Spark. This is just a simple example of how we can use the Spark shell.

SparkConf

SparkConf allows us to configure a Spark application. It sets various Spark parameters as key-value pairs, and so will usually create a SparkConf object with a SparkConf() constructor, which would then load values from the spark.* underlying Java system.

There are a few useful functions; for example, we can use the sets() function to set the configuration property. We can use the setMaster() function to set the master URL to connect to. We can use the setAppName() function to set the application name, and setSparkHome() in order to set the path where Spark will be installed on worker nodes.

You can learn more about SparkConf at https://spark.apache.org/docs/0.9.0/api/pyspark/pysaprk.conf.SparkConf-class.html.

Hands-On Big Data Analytics with PySpark

By : Rudy Lai, Bartłomiej Potaczek

Hands-On Big Data Analytics with PySpark

By: Rudy Lai, Bartłomiej Potaczek

Overview of this book

Related Content you might be interested in

Current Title:

Hands-On Big Data Analytics with PySpark

Apache Spark Quick Start Guide

Scala and Spark for Big Data Analytics

Apache Spark 2.x for Java Developers