Python Machine Learning By Example

By : Yuxi (Hayden) Liu

Python Machine Learning By Example

By: Yuxi (Hayden) Liu

Overview of this book

Data science and machine learning are some of the top buzzwords in the technical world today. A resurging interest in machine learning is due to the same factors that have made data mining and Bayesian analysis more popular than ever. This book is your entry point to machine learning. This book starts with an introduction to machine learning and the Python language and shows you how to complete the setup. Moving ahead, you will learn all the important concepts such as, exploratory data analysis, data preprocessing, feature extraction, data visualization and clustering, classification, regression and model performance evaluation. With the help of various projects included, you will find it intriguing to acquire the mechanics of several important machine learning algorithms – they are no more obscure as they thought. Also, you will be guided step by step to build your own models from scratch. Toward the end, you will gather a broad picture of the machine learning ecosystem and best practices of applying machine learning techniques. Through this book, you will learn to tackle data-driven problems and implement your solutions with the powerful yet simple language, Python. Interesting and easy-to-follow examples, to name some, news topic classification, spam email detection, online ad click-through prediction, stock prices forecast, will keep you glued till you reach your goal.

Preface

What this book covers

What you need for this book

Free Chapter

Getting Started with Python and Machine Learning

What is machine learning and why do we need it?

A very high level overview of machine learning

A brief history of the development of machine learning algorithms

Generalizing with data

Overfitting, underfitting and the bias-variance tradeoff

Avoid overfitting with feature selection and dimensionality reduction

Preprocessing, exploration, and feature engineering

Combining models

Installing software and setting up

Troubleshooting and asking for help

Summary

Exploring the 20 Newsgroups Dataset with Text Analysis Algorithms

What is NLP?

Touring powerful NLP libraries in Python

The newsgroups data

Getting the data

Thinking about features

Summary

Spam Email Detection with Naive Bayes

Getting started with classification

Types of classification

Applications of text classification

Exploring naive Bayes

Bayes' theorem by examples

The mechanics of naive Bayes

The naive Bayes implementations

Classifier performance evaluation

Model tuning and cross-validation

Summary

News Topic Classification with Support Vector Machine

Recap and inverse document frequency

Support vector machine

News topic classification with support vector machine

More examples - fetal state classification on cardiotocography with SVM

Summary

Click-Through Prediction with Tree-Based Algorithms

Brief overview of advertising click-through prediction

Getting started with two types of data, numerical and categorical

Decision tree classifier

Click-through prediction with decision tree

Random forest - feature bagging of decision tree

Summary

Click-Through Prediction with Logistic Regression

One-hot encoding - converting categorical features to numerical

Logistic regression classifier

Click-through prediction with logistic regression by gradient descent

Feature selection via random forest

Summary

Stock Price Prediction with Regression Algorithms

Brief overview of the stock market and stock price

What is regression?

Predicting stock price with regression algorithms

Summary

Best Practices

Machine learning workflow

Best practices in the data preparation stage

Best practices in the training sets generation stage

Best practices in the model training, evaluation, and selection stage

Best practices in the deployment and monitoring stage

Summary

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Combining models

In (high) school we sit together with other students, and learn together, but we are not supposed to work together during the exam. The reason is, of course, that teachers want to know what we have learned, and if we just copy exam answers from friends, we may have not learned anything. Later in life we discover that teamwork is important. For example, this book is the product of a whole team, or possibly a group of teams.

Clearly a team can produce better results than a single person. However, this goes against Occam's razor, since a single person can come up with simpler theories compared to what a team will produce. In machine learning we nevertheless prefer to have our models cooperate with the following schemes:

Bagging
Boosting
Stacking
Blending
Voting and averaging

Bagging

Bootstrap aggregating or bagging is an algorithm introduced by Leo Breiman in 1994, which applies Bootstrapping to machine learning problems. Bootstrapping is a statistical procedure, which creates datasets from existing data by sampling with replacement. Bootstrapping can be used to analyze the possible values that arithmetic mean, variance, or another quantity can assume.

The algorithm aims to reduce the chance of overfitting with the following steps:

We generate new training sets from input train data by sampling with replacement.
Fit models to each generated training set.
Combine the results of the models by averaging or majority voting.

Boosting

In the context of supervised learning we define weak learners as learners that are just a little better than a baseline such as randomly assigning classes or average values. Although weak learners are weak individually like ants, together they can do amazing things just like ants can. It makes sense to take into account the strength of each individual learner using weights. This general idea is called boosting. There are many boosting algorithms; boosting algorithms differ mostly in their weighting scheme. If you have studied for an exam, you may have applied a similar technique by identifying the type of practice questions you had trouble with and focusing on the hard problems.

Face detection in images is based on a specialized framework, which also uses boosting. Detecting faces in images or videos is a supervised learning. We give the learner examples of regions containing faces. There is an imbalance, since we usually have far more regions (about ten thousand times more) that don't have faces. A cascade of classifiers progressively filters out negative image areas stage by stage. In each progressive stage, the classifiers use progressively more features on fewer image windows. The idea is to spend the most time on image patches, which contain faces. In this context, boosting is used to select features and combine results.

Stacking

Stacking takes the outputs of machine learning estimators and then uses those as inputs for another algorithm. You can, of course, feed the output of the higher-level algorithm to another predictor. It is possible to use any arbitrary topology, but for practical reasons you should try a simple setup first as also dictated by Occam's razor.

Blending

Blending was introduced by the winners of the one million dollar Netflix prize. Netflix organized a contest with the challenge of finding the best model to recommend movies to their users. Netflix users can rate a movie with a rating of one to five stars. Obviously each user wasn't able to rate each movie, so the user movie matrix is sparse. Netflix published an anonymized training and test set. Later researchers found a way to correlate the Netflix data to IMDB data. For privacy reasons, the Netflix data is no longer available. The competition was won in 2008 by a group of teams combining their models. Blending is a form of stacking. The final estimator in blending, however, trains only on a small portion of the train data.

Voting and averaging

We can arrive at our final answer through majority voting or averaging. It's also possible to assign different weights to each model in the ensemble. For averaging, we can also use the geometric mean or the harmonic mean instead of the arithmetic mean. Usually combining the results of models, which are highly correlated to each other doesn't lead to spectacular improvements. It's better to somehow diversify the models, by using different features or different algorithms. If we find that two models are strongly correlated, we may, for example, decide to remove one of them from the ensemble, and increase the weight of the other model proportionally.

Python Machine Learning By Example

By : Yuxi (Hayden) Liu

Python Machine Learning By Example

By: Yuxi (Hayden) Liu

Overview of this book

Related Content you might be interested in

Current Title:

Python Machine Learning By Example

Mastering Machine Learning with scikit-learn

scikit-learn Cookbook

Python Machine Learning, Second Edition