Apache Spark 2 for Beginners

Apache Spark 2 for Beginners

By : Rajanarayanan Thottuvaikkatumana

Buy this Book

Apache Spark 2 for Beginners

By: Rajanarayanan Thottuvaikkatumana

Buy this Book

Overview of this book

Spark is one of the most widely-used large-scale data processing engines and runs extremely fast. It is a framework that has tools that are equally useful for application developers as well as data scientists. This book starts with the fundamentals of Spark 2 and covers the core data processing framework and API, installation, and application development setup. Then the Spark programming model is introduced through real-world examples followed by Spark SQL programming with DataFrames. An introduction to SparkR is covered next. Later, we cover the charting and plotting features of Python in conjunction with Spark data processing. After that, we take a look at Spark's stream processing, machine learning, and graph processing libraries. The last chapter combines all the skills you learned from the preceding chapters to develop a real-world Spark application. By the end of this book, you will have all the knowledge you need to develop efficient large-scale applications using Apache Spark.

Apache Spark 2 for Beginners

Credits

About the Author

About the Reviewer

www.PacktPub.com

Preface

Free Chapter

Spark Fundamentals

An overview of Apache Hadoop

Understanding Apache Spark

Installing Spark on your machines

References

Summary

Spark Programming Model

Functional programming with Spark

Understanding Spark RDD

Data transformations and actions with RDDs

Monitoring with Spark

The basics of programming with Spark

Creating RDDs from files

Understanding the Spark library stack

Reference

Summary

Spark SQL

Understanding the structure of data

Why Spark SQL?

Anatomy of Spark SQL

DataFrame programming

Understanding Aggregations in Spark SQL

Understanding multi-datasource joining with SparkSQL

Introducing datasets

Understanding Data Catalogs

References

Summary

Spark Programming with R

The need for SparkR

Basics of the R language

DataFrames in R and Spark

Spark DataFrame programming with R

Understanding aggregations in Spark R

Understanding multi-datasource joins with SparkR

References

Summary

Spark Data Analysis with Python

Charting and plotting libraries

Setting up a dataset

Data analysis use cases

Charts and plots

References

Summary

Spark Stream Processing

Data stream processing

Micro batch data processing

A log event processor

Windowed data processing

More processing options

Kafka stream processing

Spark Streaming jobs in production

References

Summary

Spark Machine Learning

Understanding machine learning

Why Spark for machine learning?

Wine quality prediction

Summary

Spark Graph Processing

Understanding graphs and their usage

The Spark GraphX library

Tennis tournament analysis

Applying the PageRank algorithm

Connected component algorithm

Understanding GraphFrames

Understanding GraphFrames queries

References

Summary

Designing Spark Applications

Lambda Architecture

Microblogging with Lambda Architecture

Implementing Lambda Architecture

Working with Spark applications

Coding style

Setting up the source code

Understanding data ingestion

Generating purposed views and queries

Understanding custom data processes

References

Summary

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Wine classification

The dataset containing various features of white wine is used in this wine quality classification use case. The following are the features of the dataset:

Fixed acidity
Volatile acidity
Citric acid
Residual sugar
Chlorides
Free sulfur dioxide
Total sulfur dioxide
Density
pH
Sulphates
Alcohol

Based on these features, the quality (score between 0 and 10) is determined. If the quality is less than 7, then it is classified as bad and a value of 0 is assigned to the label. If the quality is 7 or above, then it is classified as good and a value of 1 is assigned to the label. In other words, the classification value is the label of this dataset. Using this dataset, a model is going to be trained and then using the trained model, testing is done and predictions are made. This is a classification problem. The Logistic Regression algorithm is used to train the model. In this machine learning application use case, it deals with modeling the relationship between a dependent variable...

Apache Spark 2 for Beginners

By : Rajanarayanan Thottuvaikkatumana

Apache Spark 2 for Beginners

By: Rajanarayanan Thottuvaikkatumana

Overview of this book

Related Content you might be interested in

Current Title:

Apache Spark 2 for Beginners

Wine classification