Machine Learning with Spark - Second Edition

By : Rajdeep Dua, Manpreet Singh Ghotra

Machine Learning with Spark - Second Edition

By: Rajdeep Dua, Manpreet Singh Ghotra

Overview of this book

This book will teach you about popular machine learning algorithms and their implementation. You will learn how various machine learning concepts are implemented in the context of Spark ML. You will start by installing Spark in a single and multinode cluster. Next you'll see how to execute Scala and Python based programs for Spark ML. Then we will take a few datasets and go deeper into clustering, classification, and regression. Toward the end, we will also cover text processing using Spark ML. Once you have learned the concepts, they can be applied to implement algorithms in either green-field implementations or to migrate existing systems to this new platform. You can migrate from Mahout or Scikit to use Spark ML. By the end of this book, you will acquire the skills to leverage Spark's features to create your own scalable machine learning applications and power a modern data-driven business.

Preface

What this book covers

What you need for this book

Free Chapter

Getting Up and Running with Spark

Installing and setting up Spark locally

Spark clusters

The Spark programming model

SchemaRDD

Spark data frame

The first step to a Spark program in Scala

The first step to a Spark program in Java

The first step to a Spark program in Python

The first step to a Spark program in R

Getting Spark running on Amazon EC2

Configuring and running Spark on Amazon Elastic Map Reduce

UI in Spark

Supported machine learning algorithms by Spark

Benefits of using Spark ML as compared to existing libraries

Spark Cluster on Google Compute Engine - DataProc

Summary

Math for Machine Learning

Linear algebra

Gradient descent

Prior, likelihood, and posterior

Calculus

Plotting

Summary

Designing a Machine Learning System

What is Machine Learning?

Introducing MovieStream

Business use cases for a machine learning system

Types of machine learning models

The components of a data-driven machine learning system

An architecture for a machine learning system

Spark MLlib

Performance improvements in Spark ML over Spark MLlib

Comparing algorithms supported by MLlib

MLlib supported methods and developer APIs

MLlib vision

MLlib versions compared

Summary

Obtaining, Processing, and Preparing Data with Spark

Accessing publicly available datasets

Exploring and visualizing your data

Processing and transforming your data

Extracting useful features from your data

Summary

Building a Recommendation Engine with Spark

Types of recommendation models

Extracting the right features from your data

Training the recommendation model

Using the recommendation model

Evaluating the performance of recommendation models

FP-Growth algorithm

Summary

Building a Classification Model with Spark

Types of classification models

Extracting the right features from your data

Training classification models

Using classification models

Improving model performance and tuning parameters

Additional features

Summary

Building a Regression Model with Spark

Types of regression models

Evaluating the performance of regression models

Extracting the right features from your data

Training and using regression models

Improving model performance and tuning parameters

Summary

Building a Clustering Model with Spark

Types of clustering models

Extracting the right features from your data

K-means - training a clustering model

K-means - evaluating the performance of clustering models

Effect of iterations on WSSSE

Bisecting KMeans

Bisecting K-means - training a clustering model

Gaussian Mixture Model

Summary

Dimensionality Reduction with Spark

Types of dimensionality reduction

Extracting the right features from your data

Training a dimensionality reduction model

Using a dimensionality reduction model

Evaluating dimensionality reduction models

Summary

Advanced Text Processing with Spark

What's so special about text data?

Extracting the right features from your data

Using a tf-idf model

Evaluating the impact of text processing

Text classification with Spark 2.0

Word2Vec models

Word2Vec with Spark ML on the 20 Newsgroups dataset

Summary

Real-Time Machine Learning with Spark Streaming

Online learning

Stream processing

Online learning with Spark Streaming

Online model evaluation

Structured Streaming

Summary

Pipeline APIs for Spark ML

Introduction to pipelines

How pipelines work

Machine learning pipeline with an example

Summary

Customer Reviews

5 star

4 star

3 star

2 star

1 star

The first step to a Spark program in R

SparkR is an R package which provides a frontend to use Apache Spark from R. In Spark 1.6.0; SparkR provides a distributed data frame on large datasets. SparkR also supports distributed machine learning using MLlib. This is something you should try out while reading machine learning chapters.

SparkR DataFrames

DataFrame is a collection of data organized into names columns that are distributed. This concept is very similar to a relational database or a data frame of R but with much better optimizations. Source of these data frames could be a CSV, a TSV, Hive tables, local R data frames, and so on.

Spark distribution can be run using the ./bin/sparkR shell.

Following on from the preceding examples, we will now write an R version. We assume that you have R (R version 3.0.2 (2013-09-25)-Frisbee Sailing), R Studio and higher installed on your system (for example, most Linux and Mac OS X systems come with Python preinstalled).

The example program is included in the sample code for this chapter, in the directory named r-spark-app, which also contains the CSV data file under the data subdirectory. The project contains a script, r-script-01.R, which is provided in the following. Make sure you change PATH to appropriate value for your environment.

Sys.setenv(SPARK_HOME = "/PATH/spark-2.0.0-bin-hadoop2.7") 
.libPaths(c(file.path(Sys.getenv("SPARK_HOME"), "R", "lib"), 
 .libPaths())) 
#load the Sparkr library 
library(SparkR) 
sc <- sparkR.init(master = "local", sparkPackages="com.databricks:spark-csv_2.10:1.3.0") 
sqlContext <- sparkRSQL.init(sc) 

user.purchase.history <- "/PATH/ml-resources/spark-ml/Chapter_01/r-spark-app/data/UserPurchaseHistory.csv" 
data <- read.df(sqlContext, user.purchase.history, "com.databricks.spark.csv", header="false") 
head(data) 
count(data) 

parseFields <- function(record) { 
  Sys.setlocale("LC_ALL", "C") # necessary for strsplit() to work correctly 
  parts <- strsplit(as.character(record), ",") 
  list(name=parts[1], product=parts[2], price=parts[3]) 
} 

parsedRDD <- SparkR:::lapply(data, parseFields) 
cache(parsedRDD) 
numPurchases <- count(parsedRDD) 

sprintf("Number of Purchases : %d", numPurchases) 
getName <- function(record){ 
  record[1] 
} 

getPrice <- function(record){ 
  record[3] 
} 

nameRDD <- SparkR:::lapply(parsedRDD, getName) 
nameRDD = collect(nameRDD) 
head(nameRDD) 

uniqueUsers <- unique(nameRDD) 
head(uniqueUsers) 

priceRDD <- SparkR:::lapply(parsedRDD, function(x) { as.numeric(x$price[1])}) 
take(priceRDD,3) 

totalRevenue <- SparkR:::reduce(priceRDD, "+") 

sprintf("Total Revenue : %.2f", s) 

products <- SparkR:::lapply(parsedRDD, function(x) { list( toString(x$product[1]), 1) }) 
take(products, 5) 
productCount <- SparkR:::reduceByKey(products, "+", 2L) 
productsCountAsKey <- SparkR:::lapply(productCount, function(x) { list( as.integer(x[2][1]), x[1][1])}) 

productCount <- count(productsCountAsKey) 
mostPopular <- toString(collect(productsCountAsKey)[[productCount]][[2]]) 
sprintf("Most Popular Product : %s", mostPopular)

Run the script with the following command on the bash terminal:

  $ Rscript r-script-01.R

Your output will be similar to the following listing:

> sprintf("Number of Purchases : %d", numPurchases)
[1] "Number of Purchases : 5"

> uniqueUsers <- unique(nameRDD)
> head(uniqueUsers)
[[1]]
[[1]]$name
[[1]]$name[[1]]
[1] "John"
[[2]]
[[2]]$name
[[2]]$name[[1]]
[1] "Jack"
[[3]]
[[3]]$name
[[3]]$name[[1]]
[1] "Jill"
[[4]]
[[4]]$name
[[4]]$name[[1]]
[1] "Bob"

> sprintf("Total Revenue : %.2f", totalRevenueNum)
[1] "Total Revenue : 39.91"

> sprintf("Most Popular Product : %s", mostPopular)
[1] "Most Popular Product : iPad Cover"

Machine Learning with Spark - Second Edition

By : Rajdeep Dua, Manpreet Singh Ghotra

Machine Learning with Spark - Second Edition

By: Rajdeep Dua, Manpreet Singh Ghotra

Overview of this book

Related Content you might be interested in

Current Title:

Machine Learning with Spark - Second Edition

Kali Linux Intrusion and Exploitation Cookbook

Apache Spark 2.x Machine Learning Cookbook

Learning Spark SQL