Spark Cookbook

Book Image

Spark Cookbook

By : Rishi Yadav

Book Image

Spark Cookbook

By: Rishi Yadav

Overview of this book

Spark Cookbook

Credits

About the Author

About the Author

About the Reviewers

About the Reviewers

www.PacktPub.com

www.PacktPub.com

Preface

Free Chapter

Getting Started with Apache Spark

Getting Started with Apache Spark

Installing Spark from binaries

Building the Spark source code with Maven

Launching Spark on Amazon EC2

Deploying on a cluster in standalone mode

Deploying on a cluster with Mesos

Deploying on a cluster with YARN

Using Tachyon as an off-heap storage layer

Developing Applications with Spark

Developing Applications with Spark

Exploring the Spark shell

Developing Spark applications in Eclipse with Maven

Developing Spark applications in Eclipse with SBT

Developing a Spark application in IntelliJ IDEA with Maven

Developing a Spark application in IntelliJ IDEA with SBT

External Data Sources

External Data Sources

Loading data from the local filesystem

Loading data from HDFS

Loading data from HDFS using a custom InputFormat

Loading data from Amazon S3

Loading data from Apache Cassandra

Loading data from relational databases

Spark SQL

Understanding the Catalyst optimizer

Creating HiveContext

Inferring schema using case classes

Programmatically specifying the schema

Loading and saving data using the Parquet format

Loading and saving data using the JSON format

Loading and saving data from relational databases

Loading and saving data from an arbitrary source

Spark Streaming

Spark Streaming

Word count using Streaming

Streaming Twitter data

Streaming using Kafka

Getting Started with Machine Learning Using MLlib

Getting Started with Machine Learning Using MLlib

Creating vectors

Creating a labeled point

Creating matrices

Calculating summary statistics

Calculating correlation

Doing hypothesis testing

Creating machine learning pipelines using ML

Supervised Learning with MLlib – Regression

Supervised Learning with MLlib – Regression

Using linear regression

Understanding cost function

Doing linear regression with lasso

Doing ridge regression

Supervised Learning with MLlib – Classification

Supervised Learning with MLlib – Classification

Doing classification using logistic regression

Doing binary classification using SVM

Doing classification using decision trees

Doing classification using Random Forests

Doing classification using Gradient Boosted Trees

Doing classification with Naïve Bayes

Unsupervised Learning with MLlib

Unsupervised Learning with MLlib

Clustering using k-means

Dimensionality reduction with principal component analysis

Dimensionality reduction with singular value decomposition

Recommender Systems

Recommender Systems

Collaborative filtering using explicit feedback

Collaborative filtering using implicit feedback

Graph Processing Using GraphX

Graph Processing Using GraphX

Fundamental operations on graphs

Finding connected components

Performing neighborhood aggregation

Optimizations and Performance Tuning

Optimizations and Performance Tuning

Optimizing memory

Using compression to improve performance

Using serialization to improve performance

Optimizing garbage collection

Optimizing the level of parallelism

Understanding the future of optimization – project Tungsten

Index

Customer Reviews

5 star

0

4 star

0

3 star

0

2 star

0

1 star

0

Deploying on a cluster with Mesos

Mesos is slowly emerging as a data center operating system to manage all compute resources across a data center. Mesos runs on any computer running the Linux operating system. Mesos is built using the same principles as Linux kernel. Let's see how we can install Mesos.

How to do it...

Mesosphere provides a binary distribution of Mesos. The most recent package for the Mesos distribution can be installed from the Mesosphere repositories by performing the following steps:

Execute Mesos on Ubuntu OS with the trusty version:

$ sudo apt-key adv --keyserver keyserver.ubuntu.com --recv E56151BF DISTRO=$(lsb_release -is | tr '[:upper:]' '[:lower:]') CODENAME=$(lsb_release -cs)
$ sudo vi /etc/apt/sources.list.d/mesosphere.list

deb http://repos.mesosphere.io/Ubuntu trusty main

Update the repositories:
```
$ sudo apt-get -y update
```
Install Mesos:
```
$ sudo apt-get -y install mesos
```
To connect Spark to Mesos to integrate Spark with Mesos, make Spark binaries available to Mesos and configure the Spark driver to connect to Mesos.

Use Spark binaries from the first recipe and upload to HDFS:

$ 
hdfs dfs
 -put spark-1.4.0-bin-hadoop2.4.tgz spark-1.4.0-bin-hadoop2.4.tgz

The master URL for single master Mesos is mesos://host:5050, and for the ZooKeeper managed Mesos cluster, it is mesos://zk://host:2181.

Set the following variables in spark-env.sh:

$ sudo vi spark-env.sh
export MESOS_NATIVE_LIBRARY=/usr/local/lib/libmesos.so
export SPARK_EXECUTOR_URI= hdfs://localhost:9000/user/hduser/spark-1.4.0-bin-hadoop2.4.tgz

Run from the Scala program:

val conf = new SparkConf().setMaster("mesos://host:5050")
val sparkContext = new SparkContext(conf)

Run from the Spark shell:
```
$ spark-shell --master mesos://host:5050
```
Note
Mesos has two run modes:
Fine-grained: In fine-grained (default) mode, every Spark task runs as a separate Mesos task
Coarse-grained: This mode will launch only one long-running Spark task on each Mesos machine
To run in the coarse-grained mode, set the spark.mesos.coarse property:
```
conf.set("spark.mesos.coarse","true")
```