Mastering Machine Learning with scikit-learn

Mastering Machine Learning with scikit-learn

By : Gavin Hackeling

Buy this Book

Mastering Machine Learning with scikit-learn

By: Gavin Hackeling

Buy this Book

Overview of this book

This book examines machine learning models including logistic regression, decision trees, and support vector machines, and applies them to common problems such as categorizing documents and classifying images. It begins with the fundamentals of machine learning, introducing you to the supervised-unsupervised spectrum, the uses of training and test data, and evaluating models. You will learn how to use generalized linear models in regression problems, as well as solve problems with text and categorical features. You will be acquainted with the use of logistic regression, regularization, and the various loss functions that are used by generalized linear models. The book will also walk you through an example project that prompts you to label the most uncertain training examples. You will also use an unsupervised Hidden Markov Model to predict stock prices. By the end of the book, you will be an expert in scikit-learn and will be well versed in machine learning.

Mastering Machine Learning with scikit-learn

Credits

About the Author

About the Reviewers

www.PacktPub.com

Preface

Free Chapter

The Fundamentals of Machine Learning

Learning from experience

Machine learning tasks

Training data and test data

Performance measures, bias, and variance

An introduction to scikit-learn

Installing scikit-learn

Installing pandas and matplotlib

Summary

Linear Regression

Simple linear regression

Evaluating the model

Multiple linear regression

Polynomial regression

Regularization

Applying linear regression

Fitting models with gradient descent

Summary

Feature Extraction and Preprocessing

Extracting features from categorical variables

Extracting features from text

Extracting features from images

Data standardization

Summary

From Linear Regression to Logistic Regression

Binary classification with logistic regression

Spam filtering

Binary classification performance metrics

Calculating the F1 measure

ROC AUC

Tuning models with grid search

Multi-class classification

Multi-label classification and problem transformation

Summary

Nonlinear Classification and Regression with Decision Trees

Decision trees

Training decision trees

Decision trees with scikit-learn

Summary

Clustering with K-Means

Clustering with the K-Means algorithm

Evaluating clusters

Image quantization

Clustering to learn features

Summary

Dimensionality Reduction with PCA

An overview of PCA

Performing Principal Component Analysis

Using PCA to visualize high-dimensional data

Face recognition with PCA

Summary

The Perceptron

Activation functions

Binary classification with the perceptron

Limitations of the perceptron

Summary

From the Perceptron to Support Vector Machines

Kernels and the kernel trick

Maximum margin classification and support vectors

Classifying characters in scikit-learn

Summary

From the Perceptron to Artificial Neural Networks

Nonlinear decision boundaries

Feedforward and feedback artificial neural networks

Approximating XOR with Multilayer perceptrons

Classifying handwritten digits

Summary

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Learning from experience

Machine learning systems are often described as learning from experience either with or without supervision from humans. In supervised learning problems, a program predicts an output for an input by learning from pairs of labeled inputs and outputs; that is, the program learns from examples of the right answers. In unsupervised learning, a program does not learn from labeled data. Instead, it attempts to discover patterns in the data. For example, assume that you have collected data describing the heights and weights of people. An example of an unsupervised learning problem is dividing the data points into groups. A program might produce groups that correspond to men and women, or children and adults.

Now assume that the data is also labeled with the person's sex. An example of a supervised learning problem is inducing a rule to predict whether a person is male or female based on his or her height and weight. We will discuss algorithms and examples of supervised and unsupervised learning in the following chapters.

Supervised learning and unsupervised learning can be thought of as occupying opposite ends of a spectrum. Some types of problems, called semi-supervised learning problems, make use of both supervised and unsupervised data; these problems are located on the spectrum between supervised and unsupervised learning. An example of semi-supervised machine learning is reinforcement learning, in which a program receives feedback for its decisions, but the feedback may not be associated with a single decision. For example, a reinforcement learning program that learns to play a side-scrolling video game such as Super Mario Bros. may receive a reward when it completes a level or exceeds a certain score, and a punishment when it loses a life. However, this supervised feedback is not associated with specific decisions to run, avoid Goombas, or pick up fire flowers. While this book will discuss semi-supervised learning, we will focus primarily on supervised and unsupervised learning, as these categories include most the common machine learning problems. In the next sections, we will review supervised and unsupervised learning in more detail.

A supervised learning program learns from labeled examples of the outputs that should be produced for an input. There are many names for the output of a machine learning program. Several disciplines converge in machine learning, and many of those disciplines use their own terminology. In this book, we will refer to the output as the response variable. Other names for response variables include dependent variables, regressands, criterion variables, measured variables, responding variables, explained variables, outcome variables, experimental variables, labels, and output variables. Similarly, the input variables have several names. In this book, we will refer to the input variables as features, and the phenomena they measure as explanatory variables. Other names for explanatory variables include predictors, regressors, controlled variables, manipulated variables, and exposure variables. Response variables and explanatory variables may take real or discrete values.

The collection of examples that comprise supervised experience is called a training set. A collection of examples that is used to assess the performance of a program is called a test set. The response variable can be thought of as the answer to the question posed by the explanatory variables. Supervised learning problems learn from a collection of answers to different questions; that is, supervised learning programs are provided with the correct answers and must learn to respond correctly to unseen, but similar, questions.

Mastering Machine Learning with scikit-learn

By : Gavin Hackeling

Mastering Machine Learning with scikit-learn

By: Gavin Hackeling

Overview of this book

Related Content you might be interested in

Current Title:

Mastering Machine Learning with scikit-learn

Learning from experience