Apache Spark 2: Data Processing and Real-Time Analytics

Apache Spark 2: Data Processing and Real-Time Analytics

By : Romeo Kienzler, Md. Rezaul Karim, Sridhar Alla, Siamak Amirghodsi, Meenakshi Rajendran, Broderick Hall, Shuen Mei

Buy this Book

Apache Spark 2: Data Processing and Real-Time Analytics

By: Romeo Kienzler, Md. Rezaul Karim, Sridhar Alla, Siamak Amirghodsi, Meenakshi Rajendran, Broderick Hall, Shuen Mei

Buy this Book

Overview of this book

Apache Spark is an in-memory, cluster-based data processing system that provides a wide range of functionalities such as big data processing, analytics, machine learning, and more. With this Learning Path, you can take your knowledge of Apache Spark to the next level by learning how to expand Spark's functionality and building your own data flow and machine learning programs on this platform. You will work with the different modules in Apache Spark, such as interactive querying with Spark SQL, using DataFrames and datasets, implementing streaming analytics with Spark Streaming, and applying machine learning and deep learning techniques on Spark using MLlib and various external tools. By the end of this elaborately designed Learning Path, you will have all the knowledge you need to master Apache Spark, and build your own big data processing and analytics pipeline quickly and without any hassle. This Learning Path includes content from the following Packt products: • Mastering Apache Spark 2.x by Romeo Kienzler • Scala and Spark for Big Data Analytics by Md. Rezaul Karim, Sridhar Alla • Apache Spark 2.x Machine Learning Cookbook by Siamak Amirghodsi, Meenakshi Rajendran, Broderick Hall, Shuen MeiCookbook

Title Page

About Packt

Contributors

Preface

Free Chapter

A First Taste and What's New in Apache Spark V2

Spark machine learning

Spark Streaming

Spark SQL

Spark graph processing

Extended ecosystem

What's new in Apache Spark V2?

Cluster design

Cluster management

Cloud-based deployments

Performance

Cloud

Summary

Apache Spark Streaming

Overview

Errors and recovery

Streaming sources

Summary

Structured Streaming

The concept of continuous applications

Windowing

Increased performance with good old friends

How transparent fault tolerance and exactly-once delivery guarantee is achieved

Example - connection to a MQTT message broker

Summary

Apache Spark MLlib

Architecture

Classification with Naive Bayes

Clustering with K-Means

Artificial neural networks

Summary

Apache SparkML

What does the new API look like?

The concept of pipelines

Model evaluation

CrossValidation and hyperparameter tuning

Winning a Kaggle competition with Apache SparkML

Summary

Apache SystemML

Why do we need just another library?

A cost-based optimizer for machine learning algorithms

Performance measurements

Apache SystemML in action

Summary

Apache Spark GraphX

Overview

Graph analytics/processing with GraphX

Summary

Spark Tuning

Monitoring Spark jobs

Spark configuration

Common mistakes in Spark app development

Optimization techniques

Summary

Testing and Debugging Spark

Testing in a distributed environment

Testing Spark applications

Debugging Spark applications

Summary

Practical Machine Learning with Spark Using Scala

Introduction

Configuring IntelliJ to work with Spark and run Spark ML sample codes

Running a sample ML code from Spark

Identifying data sources for practical machine learning

Running your first program using Apache Spark 2.0 with the IntelliJ IDE

How to add graphics to your Spark program

Spark's Three Data Musketeers for Machine Learning - Perfect Together

Introduction

Creating RDDs with Spark 2.0 using internal data sources

Creating RDDs with Spark 2.0 using external data sources

Transforming RDDs with Spark 2.0 using the filter() API

Transforming RDDs with the super useful flatMap() API

Transforming RDDs with set operation APIs

RDD transformation/aggregation with groupBy() and reduceByKey()

Transforming RDDs with the zip() API

Join transformation with paired key-value RDDs

Reduce and grouping transformation with paired key-value RDDs

Creating DataFrames from Scala data structures

Operating on DataFrames programmatically without SQL

Loading DataFrames and setup from an external source

Using DataFrames with standard SQL language - SparkSQL

Working with the Dataset API using a Scala Sequence

Creating and using Datasets from RDDs and back again

Working with JSON using the Dataset API and SQL together

Functional programming with the Dataset API using domain objects

Common Recipes for Implementing a Robust Machine Learning System

Introduction

Spark's basic statistical API to help you build your own algorithms

ML pipelines for real-life machine learning applications

Normalizing data with Spark

Splitting data for training and testing

Common operations with the new Dataset API

Creating and using RDD versus DataFrame versus Dataset from a text file in Spark 2.0

LabeledPoint data structure for Spark ML

Getting access to Spark cluster in Spark 2.0

Getting access to Spark cluster pre-Spark 2.0

Getting access to SparkContext vis-a-vis SparkSession object in Spark 2.0

New model export and PMML markup in Spark 2.0

Regression model evaluation using Spark 2.0

Binary classification model evaluation using Spark 2.0

Multiclass classification model evaluation using Spark 2.0

Multilabel classification model evaluation using Spark 2.0

Using the Scala Breeze library to do graphics in Spark 2.0

Recommendation Engine that Scales with Spark

Introduction

Setting up the required data for a scalable recommendation engine in Spark 2.0

Exploring the movies data details for the recommendation system in Spark 2.0

Exploring the ratings data details for the recommendation system in Spark 2.0

Building a scalable recommendation engine using collaborative filtering in Spark 2.0

Unsupervised Clustering with Apache Spark 2.0

Introduction

Building a KMeans classifying system in Spark 2.0

Bisecting KMeans, the new kid on the block in Spark 2.0

Using Gaussian Mixture and Expectation Maximization (EM) in Spark to classify data

Classifying the vertices of a graph using Power Iteration Clustering (PIC) in Spark 2.0

Latent Dirichlet Allocation (LDA) to classify documents and text into topics

Streaming KMeans to classify data in near real-time

Implementing Text Analytics with Spark 2.0 ML Library

Introduction

Doing term frequency with Spark - everything that counts

Downloading a complete dump of Wikipedia for a real-life Spark ML project

Using Latent Semantic Analysis for text analytics with Spark 2.0

Topic modeling with Latent Dirichlet allocation in Spark 2.0

Spark Streaming and Machine Learning Library

Introduction

Structured streaming for near real-time machine learning

Streaming DataFrames for real-time machine learning

Streaming Datasets for real-time machine learning

Streaming data and debugging with queueStream

Downloading and understanding the famous Iris data for unsupervised classification

Streaming KMeans for a real-time on-line classifier

Downloading wine quality data for streaming regression

Streaming linear regression for a real-time regression

Downloading Pima Diabetes data for supervised classification

Streaming logistic regression for an on-line classifier

Other Books You May Enjoy

Leave a review - let other readers know what you think

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Contributors

About the Authors

Romeo Keinzler works as the chief data scientist in the IBM Watson IoT worldwide team, helping clients to apply advanced machine learning at scale on their IoT sensor data. He holds a Master's degree in computer science from the Swiss Federal Institute of Technology, Zurich, with a specialization in information systems, bioinformatics, and applied statistics. His current research focus is on scalable machine learning on Apache Spark. He is a contributor to various open source projects and works as an associate professor for artificial intelligence at Swiss University of Applied Sciences, Berne. He is a member of the IBM Technical Expert Council and the IBM Academy of Technology, IBM's leading brains trust.

Md. Rezaul Karim is a Research Scientist at Fraunhofer FIT, Germany. He is also a PhD candidate at RWTH Aachen University, Aachen, Germany. He holds a BSc and an MSc degree in Computer Science. Before joining Fraunhofer FIT, he worked as a Researcher at Insight Centre for Data Analytics, Ireland. Before this, he worked as a Lead Engineer at Samsung Electronics' distributed R&D Institutes in Korea, India, Turkey, and Bangladesh. Previously, he worked as a Research Assistant at the database lab, Kyung Hee University, Korea. He also worked as an R&D engineer with BMTech21 Worldwide, Korea. Before this, he worked as a Software Engineer with i2SoftTechnology, Dhaka, Bangladesh. He has more than 8 years' experience in the area of research and development with a solid understanding of algorithms and data structures in C, C++, Java, Scala, R, and Python. He has published several books, articles, and research papers concerning big data and virtualization technologies, such as Spark, Kafka, DC/OS, Docker, Mesos, Zeppelin, Hadoop, and MapReduce. He is also equally competent with deep learning technologies such as TensorFlow, DeepLearning4j, and H2O. His research interests include machine learning, deep learning, the semantic web, linked data, big data, and bioinformatics. Also he is the author of the following book titles: Large-Scale Machine Learning with Spark (Packt Publishing Ltd.) Deep Learning with TensorFlow (Packt Publishing Ltd.) Scala and Spark for Big Data Analytics (Packt Publishing Ltd.)

Sridhar Alla is a big data expert helping companies solve complex problems in distributed computing, large-scale data science and analytics practice. He presents regularly at several prestigious conferences and provides training and consulting to companies. He holds a bachelor's in computer science from JNTU, India. He loves writing code in Python, Scala, and Java. He also has extensive hands-on knowledge of several Hadoop-based technologies, TensorFlow, NoSQL, IoT, and deep learning.

Siamak Amirghodsi (Sammy) is a world-class senior technology executive leader with an entrepreneurial track record of overseeing big data strategies, cloud transformation, quantitative risk management, advanced analytics, large-scale regulatory data platforming, enterprise architecture, technology road mapping, multi-project execution, and organizational streamlining in Fortune 20 environments in a global setting. Siamak is a hands-on big data, cloud, machine learning, and AI expert, and is currently overseeing the large-scale cloud data platforming and advanced risk analytics build out for a tier-1 financial institution in the United States. Siamak's interests include building advanced technical teams, executive management, Spark, Hadoop, big data analytics, AI, deep learning nets, TensorFlow, cognitive models, swarm algorithms, real-time streaming systems, quantum computing, financial risk management, trading signal discovery, econometrics, long-term financial cycles, IoT, blockchain, probabilistic graphical models, cryptography, and NLP.

Meenakshi Rajendran is a hands-on big data analytics and data governance manager with expertise in large-scale data platforming and machine learning program execution on a global scale. She is experienced in the end-to-end delivery of data analytics and data science products for leading financial institutions. Meenakshi holds a master's degree in business administration and is a certified PMP with over 13 years of experience in global software delivery environments. She not only understands the underpinnings of big data and data science technology but also has a solid understanding of the human side of the equation as well.Meenakshi’s favorite languages are Python, R, Julia, and Scala. Her areas of research and interest are Apache Spark, cloud, regulatory data governance, machine learning, Cassandra, and managing global data teams at scale. In her free time, she dabbles in software engineering management literature, cognitive psychology, and chess for relaxation.

Broderick Hall is a hands-on big data analytics expert and holds a master’s degree in computer science with 20 years of experience in designing and developing complex enterprise-wide software applications with real-time and regulatory requirements at a global scale. He has an extensive experience in designing and building real-time financial applications for some of the largest financial institutions and exchanges in USA. He is a deep learning early adopter and is currently working on a large-scale cloud-based data platform with deep learning net augmentation.Shuen Mei is a big data analytic platforms expert with 15+ years of experience in the financial services industry. He is experienced in designing, building, and executing large-scale, enterprise-distributed financial systems with mission-critical low-latency requirements. He is certified in the Apache Spark, Cloudera Big Data platform, including Developer, Admin, and HBase.Shuen is also a certified AWS solutions architect with emphasis on peta-byte range real-time data platform systems. Shuen is a skilled software engineer with extensive experience in delivering infrastructure, code, data architecture, and performance tuning solutions in trading and finance for Fortune 100 companies.

Packt Is Searching for Authors Like You

If you're interested in becoming an author for Packt, please visit authors.packtpub.com and apply today. We have worked with thousands of developers and tech professionals, just like you, to help them share their insight with the global tech community. You can make a general application, apply for a specific hot topic that we are recruiting an author for, or submit your own idea.

Apache Spark 2: Data Processing and Real-Time Analytics

By : Romeo Kienzler, Md. Rezaul Karim, Sridhar Alla, Siamak Amirghodsi, Meenakshi Rajendran, Broderick Hall, Shuen Mei

Apache Spark 2: Data Processing and Real-Time Analytics

By: Romeo Kienzler, Md. Rezaul Karim, Sridhar Alla, Siamak Amirghodsi, Meenakshi Rajendran, Broderick Hall, Shuen Mei

Overview of this book

Related Content you might be interested in

Current Title:

Apache Spark 2: Data Processing and Real-Time Analytics

Contributors

About the Authors

Packt Is Searching for Authors Like You