scikit-learn Cookbook - Second Edition

By : Trent Hauck

scikit-learn Cookbook - Second Edition

By: Trent Hauck

Overview of this book

Python is quickly becoming the go-to language for analysts and data scientists due to its simplicity and flexibility, and within the Python data space, scikit-learn is the unequivocal choice for machine learning. This book includes walk throughs and solutions to the common as well as the not-so-common problems in machine learning, and how scikit-learn can be leveraged to perform various machine learning tasks effectively. The second edition begins with taking you through recipes on evaluating the statistical properties of data and generates synthetic data for machine learning modelling. As you progress through the chapters, you will comes across recipes that will teach you to implement techniques like data pre-processing, linear regression, logistic regression, K-NN, Naïve Bayes, classification, decision trees, Ensembles and much more. Furthermore, you’ll learn to optimize your models with multi-class classification, cross validation, model evaluation and dive deeper in to implementing deep learning with scikit-learn. Along with covering the enhanced features on model section, API and new features like classifiers, regressors and estimators the book also contains recipes on evaluating and fine-tuning the performance of your model. By the end of this book, you will have explored plethora of features offered by scikit-learn for Python to solve any machine learning problem you come across.

Preface

What this book covers

Who this book is for

What you need for this book

Conventions

Reader feedback

Customer support

Free Chapter

High-Performance Machine Learning – NumPy

Introduction

NumPy basics

Loading the iris dataset

Viewing the iris dataset

Viewing the iris dataset with Pandas

Plotting with NumPy and matplotlib

A minimal machine learning recipe – SVM classification

Introducing cross-validation

Putting it all together

Machine learning overview – classification versus regression

Pre-Model Workflow and Pre-Processing

Introduction

Creating sample data for toy analysis

Scaling data to the standard normal distribution

Creating binary features through thresholding

Working with categorical variables

Imputing missing values through various strategies

A linear model in the presence of outliers

Putting it all together with pipelines

Using Gaussian processes for regression

Using SGD for regression

Dimensionality Reduction

Introduction

Reducing dimensionality with PCA

Using factor analysis for decomposition

Using kernel PCA for nonlinear dimensionality reduction

Using truncated SVD to reduce dimensionality

Using decomposition to classify with DictionaryLearning

Doing dimensionality reduction with manifolds – t-SNE

Testing methods to reduce dimensionality with pipelines

Linear Models with scikit-learn

Introduction

Fitting a line through data

Fitting a line through data with machine learning

Evaluating the linear regression model

Using ridge regression to overcome linear regression's shortfalls

Optimizing the ridge regression parameter

Using sparsity to regularize models

Taking a more fundamental approach to regularization with LARS

References

Linear Models – Logistic Regression

Introduction

Loading data from the UCI repository

Viewing the Pima Indians diabetes dataset with pandas

Looking at the UCI Pima Indians dataset web page

Machine learning with logistic regression

Examining logistic regression errors with a confusion matrix

Varying the classification threshold in logistic regression

Receiver operating characteristic – ROC analysis

Plotting an ROC curve without context

Putting it all together – UCI breast cancer dataset

Building Models with Distance Metrics

Introduction

Using k-means to cluster data

Optimizing the number of centroids

Assessing cluster correctness

Using MiniBatch k-means to handle more data

Quantizing an image with k-means clustering

Finding the closest object in the feature space

Probabilistic clustering with Gaussian mixture models

Using k-means for outlier detection

Using KNN for regression

Cross-Validation and Post-Model Workflow

Introduction

Selecting a model with cross-validation

K-fold cross validation

Balanced cross-validation

Cross-validation with ShuffleSplit

Time series cross-validation

Grid search with scikit-learn

Randomized search with scikit-learn

Classification metrics

Regression metrics

Clustering metrics

Using dummy estimators to compare results

Feature selection

Feature selection on L1 norms

Persisting models with joblib or pickle

Support Vector Machines

Introduction

Classifying data with a linear SVM

Optimizing an SVM

Multiclass classification with SVM

Support vector regression

Tree Algorithms and Ensembles

Introduction

Doing basic classifications with decision trees

Visualizing a decision tree with pydot

Tuning a decision tree

Using decision trees for regression

Reducing overfitting with cross-validation

Implementing random forest regression

Bagging regression with nearest neighbors

Tuning gradient boosting trees

Tuning an AdaBoost regressor

Writing a stacking aggregator with scikit-learn

Text and Multiclass Classification with scikit-learn

Using LDA for classification

Working with QDA – a nonlinear LDA

Using SGD for classification

Classifying documents with Naive Bayes

Label propagation with semi-supervised learning

Neural Networks

Introduction

Perceptron classifier

Neural network – multilayer perceptron

Stacking with a neural network

Create a Simple Estimator

Introduction

Create a simple estimator

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Viewing the iris dataset

Now that we've loaded the dataset, let's examine what is in it. The iris dataset pertains to a supervised classification problem.

How to do it...

To access the observation variables, type:

iris.data

This outputs a NumPy array:

array([[ 5.1,  3.5,  1.4,  0.2],
       [ 4.9,  3. ,  1.4,  0.2],
       [ 4.7,  3.2,  1.3,  0.2], 
#...rest of output suppressed because of length

Let's examine the NumPy array:

iris.data.shape

This returns:

(150L, 4L)

This means that the data is 150 rows by 4 columns. Let's look at the first row:

iris.data[0]

array([ 5.1,  3.5,  1.4,  0.2])

The NumPy array for the first row has four numbers.

To determine what they mean, type:

iris.feature_names
['sepal length (cm)',
 'sepal width (cm)',
 'petal length (cm)',
 'petal width (cm)']

The feature or column names name the data. They are strings, and in this case, they correspond to dimensions in different types of flowers. Putting it all together, we have 150 examples of flowers with four measurements per flower in centimeters. For example, the first flower has measurements of 5.1 cm for sepal length, 3.5 cm for sepal width, 1.4 cm for petal length, and 0.2 cm for petal width. Now, let's look at the output variable in a similar manner:

iris.target

This yields an array of outputs: 0, 1, and 2. There are only three outputs. Type this:

iris.target.shape

You get a shape of:

(150L,)

This refers to an array of length 150 (150 x 1). Let's look at what the numbers refer to:

iris.target_names

array(['setosa', 'versicolor', 'virginica'], 
      dtype='|S10')

The output of the iris.target_names variable gives the English names for the numbers in the iris.target variable. The number zero corresponds to the setosa flower, number one corresponds to the versicolor flower, and number two corresponds to the virginica flower. Look at the first row of iris.target:

iris.target[0]

This produces zero, and thus the first row of observations we examined before correspond to the setosa flower.

How it works...

In machine learning, we often deal with data tables and two-dimensional arrays corresponding to examples. In the iris set, we have 150 observations of flowers of three types. With new observations, we would like to predict which type of flower those observations correspond to. The observations in this case are measurements in centimeters. It is important to look at the data pertaining to real objects. Quoting my high school physics teacher, "Do not forget the units!"

The iris dataset is intended to be for a supervised machine learning task because it has a target array, which is the variable we desire to predict from the observation variables. Additionally, it is a classification problem, as there are three numbers we can predict from the observations, one for each type of flower. In a classification problem, we are trying to distinguish between categories. The simplest case is binary classification. The iris dataset, with three flower categories, is a multi-class classification problem.

There's more...

With the same data, we can rephrase the problem in many ways, or formulate new problems. What if we want to determine relationships between the observations? We can define the petal width as the target variable. We can rephrase the problem as a regression problem and try to predict the target variable as a real number, not just three categories. Fundamentally, it comes down to what we intend to predict. Here, we desire to predict a type of flower.

scikit-learn Cookbook - Second Edition

By : Trent Hauck

scikit-learn Cookbook - Second Edition

By: Trent Hauck

Overview of this book

Related Content you might be interested in

Current Title:

scikit-learn Cookbook - Second Edition

Hands-On Machine Learning with scikit-learn and Scientific Python Toolkits

Machine Learning with scikit-learn Quick Start Guide

Hands-On Ensemble Learning with Python