Python Machine Learning, Second Edition - Second Edition

Book Image

Python Machine Learning, Second Edition - Second Edition

By : Sebastian Raschka, Vahid Mirjalili

Book Image

Python Machine Learning, Second Edition - Second Edition

By: Sebastian Raschka, Vahid Mirjalili

Overview of this book

Publisher's Note: This edition from 2017 is outdated and is not compatible with TensorFlow 2 or any of the most recent updates to Python libraries. A new third edition, updated for 2020 and featuring TensorFlow 2 and the latest in scikit-learn, reinforcement learning, and GANs, has now been published. Machine learning is eating the software world, and now deep learning is extending machine learning. Understand and work at the cutting edge of machine learning, neural networks, and deep learning with this second edition of Sebastian Raschka’s bestselling book, Python Machine Learning. Using Python's open source libraries, this book offers the practical knowledge and techniques you need to create and contribute to machine learning, deep learning, and modern data analysis. Fully extended and modernized, Python Machine Learning Second Edition now includes the popular TensorFlow 1.x deep learning library. The scikit-learn code has also been fully updated to v0.18.1 to include improvements and additions to this versatile machine learning library. Sebastian Raschka and Vahid Mirjalili’s unique insight and expertise introduce you to machine learning and deep learning algorithms from scratch, and show you how to apply them to practical industry challenges using realistic and interesting examples. By the end of the book, you’ll be ready to meet the new data analysis opportunities. If you’ve read the first edition of this book, you’ll be delighted to find a balance of classical ideas and modern insights into machine learning. Every chapter has been critically updated, and there are new chapters on key technologies. You’ll be able to learn and work with TensorFlow 1.x more deeply than ever before, and get essential coverage of the Keras neural network library, along with updates to scikit-learn 0.18.1.

Python Machine Learning Second Edition

Python Machine Learning Second Edition

Credits

About the Authors

About the Authors

About the Reviewers

About the Reviewers

www.PacktPub.com

www.PacktPub.com

Packt is Searching for Authors Like You

Packt is Searching for Authors Like You

Preface

Free Chapter

Giving Computers the Ability to Learn from Data

Giving Computers the Ability to Learn from Data

Building intelligent machines to transform data into knowledge

The three different types of machine learning

Introduction to the basic terminology and notations

A roadmap for building machine learning systems

Using Python for machine learning

Training Simple Machine Learning Algorithms for Classification

Training Simple Machine Learning Algorithms for Classification

Artificial neurons – a brief glimpse into the early history of machine learning

Implementing a perceptron learning algorithm in Python

Adaptive linear neurons and the convergence of learning

A Tour of Machine Learning Classifiers Using scikit-learn

A Tour of Machine Learning Classifiers Using scikit-learn

Choosing a classification algorithm

First steps with scikit-learn – training a perceptron

Modeling class probabilities via logistic regression

Maximum margin classification with support vector machines

Solving nonlinear problems using a kernel SVM

Decision tree learning

K-nearest neighbors – a lazy learning algorithm

Building Good Training Sets – Data Preprocessing

Building Good Training Sets – Data Preprocessing

Dealing with missing data

Handling categorical data

Partitioning a dataset into separate training and test sets

Bringing features onto the same scale

Selecting meaningful features

Assessing feature importance with random forests

Compressing Data via Dimensionality Reduction

Compressing Data via Dimensionality Reduction

Unsupervised dimensionality reduction via principal component analysis

Supervised data compression via linear discriminant analysis

Using kernel principal component analysis for nonlinear mappings

Learning Best Practices for Model Evaluation and Hyperparameter Tuning

Learning Best Practices for Model Evaluation and Hyperparameter Tuning

Streamlining workflows with pipelines

Using k-fold cross-validation to assess model performance

Debugging algorithms with learning and validation curves

Fine-tuning machine learning models via grid search

Looking at different performance evaluation metrics

Dealing with class imbalance

Combining Different Models for Ensemble Learning

Combining Different Models for Ensemble Learning

Learning with ensembles

Combining classifiers via majority vote

Bagging – building an ensemble of classifiers from bootstrap samples

Leveraging weak learners via adaptive boosting

Applying Machine Learning to Sentiment Analysis

Applying Machine Learning to Sentiment Analysis

Preparing the IMDb movie review data for text processing

Introducing the bag-of-words model

Training a logistic regression model for document classification

Working with bigger data – online algorithms and out-of-core learning

Topic modeling with Latent Dirichlet Allocation

Embedding a Machine Learning Model into a Web Application

Embedding a Machine Learning Model into a Web Application

Serializing fitted scikit-learn estimators

Setting up an SQLite database for data storage

Developing a web application with Flask

Turning the movie review classifier into a web application

Deploying the web application to a public server

Predicting Continuous Target Variables with Regression Analysis

Predicting Continuous Target Variables with Regression Analysis

Introducing linear regression

Exploring the Housing dataset

Implementing an ordinary least squares linear regression model

Fitting a robust regression model using RANSAC

Evaluating the performance of linear regression models

Using regularized methods for regression

Turning a linear regression model into a curve – polynomial regression

Dealing with nonlinear relationships using random forests

Working with Unlabeled Data – Clustering Analysis

Working with Unlabeled Data – Clustering Analysis

Grouping objects by similarity using k-means

Organizing clusters as a hierarchical tree

Locating regions of high density via DBSCAN

Implementing a Multilayer Artificial Neural Network from Scratch

Implementing a Multilayer Artificial Neural Network from Scratch

Modeling complex functions with artificial neural networks

Classifying handwritten digits

Training an artificial neural network

About the convergence in neural networks

A few last words about the neural network implementation

Parallelizing Neural Network Training with TensorFlow

Parallelizing Neural Network Training with TensorFlow

TensorFlow and training performance

Training neural networks efficiently with high-level TensorFlow APIs

Choosing activation functions for multilayer networks

Going Deeper – The Mechanics of TensorFlow

Going Deeper – The Mechanics of TensorFlow

Key features of TensorFlow

TensorFlow ranks and tensors

Understanding TensorFlow's computation graphs

Placeholders in TensorFlow

Variables in TensorFlow

Building a regression model

Executing objects in a TensorFlow graph using their names

Saving and restoring a model in TensorFlow

Transforming Tensors as multidimensional data arrays

Utilizing control flow mechanics in building graphs

Visualizing the graph with TensorBoard

Classifying Images with Deep Convolutional Neural Networks

Classifying Images with Deep Convolutional Neural Networks

Building blocks of convolutional neural networks

Putting everything together to build a CNN

Implementing a deep convolutional neural network using TensorFlow

Modeling Sequential Data Using Recurrent Neural Networks

Modeling Sequential Data Using Recurrent Neural Networks

Introducing sequential data

RNNs for modeling sequences

Implementing a multilayer RNN for sequence modeling in TensorFlow

Project one – performing sentiment analysis of IMDb movie reviews using multilayer RNNs

Project two – implementing an RNN for character-level language modeling in TensorFlow

Chapter and book summary

Index

Customer Reviews

5 star

0

4 star

0

3 star

0

2 star

0

1 star

0

Assessing feature importance with random forests

In previous sections, you learned how to use L1 regularization to zero out irrelevant features via logistic regression, and use the SBS algorithm for feature selection and apply it to a KNN algorithm. Another useful approach to select relevant features from a dataset is to use a random forest, an ensemble technique that we introduced in Chapter 3, A Tour of Machine Learning Classifiers Using scikit-learn. Using a random forest, we can measure the feature importance as the averaged impurity decrease computed from all decision trees in the forest, without making any assumptions about whether our data is linearly separable or not. Conveniently, the random forest implementation in scikit-learn already collects the feature importance values for us so that we can access them via the feature_importances_ attribute after fitting a RandomForestClassifier. By executing the following code, we will now train a forest of 500 trees on the Wine dataset and...