Machine Learning with Spark

Machine Learning with Spark

By : Nick Pentreath

Buy this Book

Machine Learning with Spark

By: Nick Pentreath

Buy this Book

Overview of this book

<p>Apache Spark is a framework for distributed computing that is designed from the ground up to be optimized for low latency tasks and in-memory data storage. It is one of the few frameworks for parallel computing that combines speed, scalability, in-memory processing, and fault tolerance with ease of programming and a flexible, expressive, and powerful API design.</p> <p>This book guides you through the basics of Spark's API used to load and process data and prepare the data to use as input to the various machine learning models. There are detailed examples and real-world use cases for you to explore common machine learning models including recommender systems, classification, regression, clustering, and dimensionality reduction. You will cover advanced topics such as working with large-scale text data, and methods for online machine learning and model evaluation using Spark Streaming.</p>

Machine Learning with Spark

Credits

About the Author

Acknowledgments

About the Reviewers

www.PacktPub.com

Preface

Free Chapter

Getting Up and Running with Spark

Installing and setting up Spark locally

Spark clusters

The Spark programming model

The first step to a Spark program in Scala

The first step to a Spark program in Java

The first step to a Spark program in Python

Getting Spark running on Amazon EC2

Summary

Designing a Machine Learning System

Introducing MovieStream

Business use cases for a machine learning system

Types of machine learning models

The components of a data-driven machine learning system

An architecture for a machine learning system

Summary

Obtaining, Processing, and Preparing Data with Spark

Accessing publicly available datasets

Exploring and visualizing your data

Processing and transforming your data

Extracting useful features from your data

Summary

Building a Recommendation Engine with Spark

Types of recommendation models

Extracting the right features from your data

Training the recommendation model

Using the recommendation model

Evaluating the performance of recommendation models

Summary

Building a Classification Model with Spark

Types of classification models

Extracting the right features from your data

Training classification models

Using classification models

Evaluating the performance of classification models

Improving model performance and tuning parameters

Summary

Building a Regression Model with Spark

Types of regression models

Extracting the right features from your data

Training and using regression models

Evaluating the performance of regression models

Improving model performance and tuning parameters

Summary

Building a Clustering Model with Spark

Types of clustering models

Extracting the right features from your data

Training a clustering model

Making predictions using a clustering model

Evaluating the performance of clustering models

Tuning parameters for clustering models

Summary

Dimensionality Reduction with Spark

Types of dimensionality reduction

Extracting the right features from your data

Training a dimensionality reduction model

Using a dimensionality reduction model

Evaluating dimensionality reduction models

Summary

Advanced Text Processing with Spark

What's so special about text data?

Extracting the right features from your data

Using a TF-IDF model

Evaluating the impact of text processing

Word2Vec models

Summary

Real-time Machine Learning with Spark Streaming

Online learning

Stream processing

Creating a Spark Streaming application

Online learning with Spark Streaming

Online model evaluation

Summary

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Improving model performance and tuning parameters

So, what went wrong? Why have our sophisticated models achieved nothing better than random chance? Is there a problem with our models?

Recall that we started out by just throwing the data at our model. In fact, we didn't even throw all our data at the model, just the numeric columns that were easy to use. Furthermore, we didn't do a lot of analysis on these numeric features.

Feature standardization

Many models that we employ make inherent assumptions about the distribution or scale of input data. One of the most common forms of assumption is about normally-distributed features. Let's take a deeper look at the distribution of our features.

To do this, we can represent the feature vectors as a distributed matrix in MLlib, using the RowMatrix class. RowMatrix is an RDD made up of vector, where each vector is a row of our matrix.

The RowMatrix class comes with some useful methods to operate on the matrix, one of which is a utility to compute statistics...

Machine Learning with Spark

By : Nick Pentreath

Machine Learning with Spark

By: Nick Pentreath

Overview of this book

Related Content you might be interested in

Current Title:

Machine Learning with Spark

Improving model performance and tuning parameters

Feature standardization