scikit-learn Cookbook - Second Edition

By : Trent Hauck

scikit-learn Cookbook - Second Edition

By: Trent Hauck

Overview of this book

Python is quickly becoming the go-to language for analysts and data scientists due to its simplicity and flexibility, and within the Python data space, scikit-learn is the unequivocal choice for machine learning. This book includes walk throughs and solutions to the common as well as the not-so-common problems in machine learning, and how scikit-learn can be leveraged to perform various machine learning tasks effectively. The second edition begins with taking you through recipes on evaluating the statistical properties of data and generates synthetic data for machine learning modelling. As you progress through the chapters, you will comes across recipes that will teach you to implement techniques like data pre-processing, linear regression, logistic regression, K-NN, Naïve Bayes, classification, decision trees, Ensembles and much more. Furthermore, you’ll learn to optimize your models with multi-class classification, cross validation, model evaluation and dive deeper in to implementing deep learning with scikit-learn. Along with covering the enhanced features on model section, API and new features like classifiers, regressors and estimators the book also contains recipes on evaluating and fine-tuning the performance of your model. By the end of this book, you will have explored plethora of features offered by scikit-learn for Python to solve any machine learning problem you come across.

Preface

What this book covers

Who this book is for

What you need for this book

Conventions

Reader feedback

Customer support

Free Chapter

High-Performance Machine Learning – NumPy

Introduction

NumPy basics

Loading the iris dataset

Viewing the iris dataset

Viewing the iris dataset with Pandas

Plotting with NumPy and matplotlib

A minimal machine learning recipe – SVM classification

Introducing cross-validation

Putting it all together

Machine learning overview – classification versus regression

Pre-Model Workflow and Pre-Processing

Introduction

Creating sample data for toy analysis

Scaling data to the standard normal distribution

Creating binary features through thresholding

Working with categorical variables

Imputing missing values through various strategies

A linear model in the presence of outliers

Putting it all together with pipelines

Using Gaussian processes for regression

Using SGD for regression

Dimensionality Reduction

Introduction

Reducing dimensionality with PCA

Using factor analysis for decomposition

Using kernel PCA for nonlinear dimensionality reduction

Using truncated SVD to reduce dimensionality

Using decomposition to classify with DictionaryLearning

Doing dimensionality reduction with manifolds – t-SNE

Testing methods to reduce dimensionality with pipelines

Linear Models with scikit-learn

Introduction

Fitting a line through data

Fitting a line through data with machine learning

Evaluating the linear regression model

Using ridge regression to overcome linear regression's shortfalls

Optimizing the ridge regression parameter

Using sparsity to regularize models

Taking a more fundamental approach to regularization with LARS

References

Linear Models – Logistic Regression

Introduction

Loading data from the UCI repository

Viewing the Pima Indians diabetes dataset with pandas

Looking at the UCI Pima Indians dataset web page

Machine learning with logistic regression

Examining logistic regression errors with a confusion matrix

Varying the classification threshold in logistic regression

Receiver operating characteristic – ROC analysis

Plotting an ROC curve without context

Putting it all together – UCI breast cancer dataset

Building Models with Distance Metrics

Introduction

Using k-means to cluster data

Optimizing the number of centroids

Assessing cluster correctness

Using MiniBatch k-means to handle more data

Quantizing an image with k-means clustering

Finding the closest object in the feature space

Probabilistic clustering with Gaussian mixture models

Using k-means for outlier detection

Using KNN for regression

Cross-Validation and Post-Model Workflow

Introduction

Selecting a model with cross-validation

K-fold cross validation

Balanced cross-validation

Cross-validation with ShuffleSplit

Time series cross-validation

Grid search with scikit-learn

Randomized search with scikit-learn

Classification metrics

Regression metrics

Clustering metrics

Using dummy estimators to compare results

Feature selection

Feature selection on L1 norms

Persisting models with joblib or pickle

Support Vector Machines

Introduction

Classifying data with a linear SVM

Optimizing an SVM

Multiclass classification with SVM

Support vector regression

Tree Algorithms and Ensembles

Introduction

Doing basic classifications with decision trees

Visualizing a decision tree with pydot

Tuning a decision tree

Using decision trees for regression

Reducing overfitting with cross-validation

Implementing random forest regression

Bagging regression with nearest neighbors

Tuning gradient boosting trees

Tuning an AdaBoost regressor

Writing a stacking aggregator with scikit-learn

Text and Multiclass Classification with scikit-learn

Using LDA for classification

Working with QDA – a nonlinear LDA

Using SGD for classification

Classifying documents with Naive Bayes

Label propagation with semi-supervised learning

Neural Networks

Introduction

Perceptron classifier

Neural network – multilayer perceptron

Stacking with a neural network

Create a Simple Estimator

Introduction

Create a simple estimator

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Viewing the iris dataset with Pandas

In this recipe we will use the handy pandas data analysis library to view and visualize the iris dataset. It contains the notion o, a dataframe which might be familiar to you if you use the language R's dataframe.

How to do it...

You can view the iris dataset with Pandas, a library built on top of NumPy:

Create a dataframe with the observation variables iris.data, and column names columns, as arguments:

import pandas as pd
iris_df = pd.DataFrame(iris.data, columns = iris.feature_names)

The dataframe is more user-friendly than the NumPy array.

Look at a quick histogram of the values in the dataframe for sepal length:

iris_df['sepal length (cm)'].hist(bins=30)

You can also color the histogram by the target variable:

for class_number in np.unique(iris.target):
    plt.figure(1)
    iris_df['sepal length (cm)'].iloc[np.where(iris.target == class_number)[0]].hist(bins=30)

Here, iterate through the target numbers for each flower and draw a color histogram for each. Consider this line:

np.where(iris.target== class_number)[0]

It finds the NumPy index location for each class of flower:

Observe that the histograms overlap. This encourages us to model the three histograms as three normal distributions. This is possible in a machine learning manner if we model the training data only as three normal distributions, not the whole set. Then we use the test set to test the three normal distribution models we just made up. Finally, we test the accuracy of our predictions on the test set.

How it works...

The dataframe data object is a 2D NumPy array with column names and row names. In data science, the fundamental data object looks like a 2D table, possibly because of SQL's long history. NumPy allows for 3D arrays, cubes, 4D arrays, and so on. These also come up often.

scikit-learn Cookbook - Second Edition

By : Trent Hauck

scikit-learn Cookbook - Second Edition

By: Trent Hauck

Overview of this book

Related Content you might be interested in

Current Title:

scikit-learn Cookbook - Second Edition

Hands-On Machine Learning with scikit-learn and Scientific Python Toolkits

Machine Learning with scikit-learn Quick Start Guide

Hands-On Ensemble Learning with Python