Machine Learning Algorithms - Second Edition

Overview of this book

Machine learning has gained tremendous popularity for its powerful and fast predictions with large datasets. However, the true forces behind its powerful output are the complex algorithms involving substantial statistical analysis that churn large datasets and generate substantial insight. This second edition of Machine Learning Algorithms walks you through prominent development outcomes that have taken place relating to machine learning algorithms, which constitute major contributions to the machine learning process and help you to strengthen and master statistical interpretation across the areas of supervised, semi-supervised, and reinforcement learning. Once the core concepts of an algorithm have been covered, you’ll explore real-world examples based on the most diffused libraries, such as scikit-learn, NLTK, TensorFlow, and Keras. You will discover new topics such as principal component analysis (PCA), independent component analysis (ICA), Bayesian regression, discriminant analysis, advanced clustering, and gaussian mixture. By the end of this book, you will have studied machine learning algorithms and be able to put them into production to make your machine learning applications more innovative.

Preface

Who this book is for

What this book covers

To get the most out of this book

Get in touch

Free Chapter

A Gentle Introduction to Machine Learning

Introduction – classic and adaptive machines

Only learning matters

Beyond machine learning – deep learning and bio-inspired adaptive systems

Machine learning and big data

Summary

Important Elements in Machine Learning

Data formats

Learnability

Introduction to statistical learning concepts

Class balancing

Elements of information theory

Summary

Feature Selection and Feature Engineering

scikit-learn toy datasets

Creating training and test sets

Managing categorical data

Managing missing features

Data scaling and normalization

Feature selection and filtering

Principal Component Analysis

Independent Component Analysis

Atom extraction and dictionary learning

Visualizing high-dimensional datasets using t-SNE

Summary

Regression Algorithms

Linear models for regression

A bidimensional example

Linear regression with scikit-learn and higher dimensionality

Ridge, Lasso, and ElasticNet

Robust regression

Bayesian regression

Polynomial regression

Isotonic regression

Summary

Linear Classification Algorithms

Linear classification

Logistic regression

Implementation and optimizations

Stochastic gradient descent algorithms

Passive-aggressive algorithms

Finding the optimal hyperparameters through a grid search

Classification metrics

ROC curve

Summary

Naive Bayes and Discriminant Analysis

Bayes' theorem

Naive Bayes classifiers

Naive Bayes in scikit-learn

Discriminant analysis

Summary

Support Vector Machines

Linear SVM

SVMs with scikit-learn

Kernel-based classification

ν-Support Vector Machines

Support Vector Regression

Introducing semi-supervised Support Vector Machines (S3VM)

Summary

Decision Trees and Ensemble Learning

Binary Decision Trees

Decision Tree classification with scikit-learn

Decision Tree regression

Introduction to Ensemble Learning

Summary

Clustering Fundamentals

Clustering basics

k-NN

Gaussian mixture

K-means

Evaluation methods based on the ground truth

Summary

Advanced Clustering

DBSCAN

Spectral Clustering

Online Clustering

Biclustering

Summary

Hierarchical Clustering

Hierarchical strategies

Agglomerative Clustering

Summary

Introducing Recommendation Systems

Naive user-based systems

Content-based systems

Model-free (or memory-based) collaborative filtering

Model-based collaborative filtering

Summary

Introducing Natural Language Processing

NLTK and built-in corpora

The Bag-of-Words strategy

Part-of-Speech

A sample text classifier based on the Reuters corpus

Summary

Topic Modeling and Sentiment Analysis in NLP

Topic modeling

Introducing Word2vec with Gensim

Sentiment analysis

Summary

Introducing Neural Networks

Deep learning at a glance

MLPs with Keras

Summary

Advanced Deep Learning Models

Deep model layers

An example of a deep convolutional network with Keras

An example of an LSTM network with Keras

A brief introduction to TensorFlow

Summary

Creating a Machine Learning Architecture

Machine learning architectures

Scikit-learn tools for machine learning architectures

Summary

Other Books You May Enjoy

Leave a review - let other readers know what you think

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Machine learning and big data

Another area that can be exploited using machine learning is big data. After the first release of Apache Hadoop, which implemented an efficient MapReduce algorithm, the amount of information managed in different business contexts grew exponentially. At the same time, the opportunity to use it for machine learning purposes arose and several applications such as mass collaborative filtering became a reality.

Imagine an online store with 1 million users and only 1,000 products. Consider a matrix where each user is associated with every product by an implicit or explicit ranking. This matrix will contain 1,000,000 x 1,000 cells, and even if the number of products is very limited, any operation performed on it will be slow and memory-consuming. Instead, using a cluster, together with parallel algorithms, such a problem disappears, and operations with a higher dimensionality can be carried out in a very short time.

Think about training an image classifier with 1 million samples. A single instance needs to iterate several times, processing small batches of pictures. Even if this problem can be performed using a streaming approach (with a limited amount of memory), it's not surprising to wait even for a few days before the model begins to perform well. Adopting a big data approach instead, it's possible to asynchronously train several local models, periodically share the updates, and re-synchronize them all with a master model. This technique has also been exploited to solve some reinforcement learning problems, where many agents (often managed by different threads) played the same game, providing their periodical contribution to a global intelligence.

Not every machine learning problem is suitable for big data, and not all big datasets are really useful when training models. However, their conjunction in particular situations can lead to extraordinary results by removing many limitations that often affect smaller scenarios. Unfortunately, both machine learning and big data are topics subject to continuous hype, hence one of the tasks that an engineer/scientist has to accomplish is understanding when a particular technology is really helpful and when its burden can be heavier than the benefits. Modern computers often have enough resources to process datasets that, a few years ago, were easily considered big data. Therefore, I invite the reader to carefully analyze each situation and think about the problem from a business viewpoint as well. A Spark cluster has a cost that is sometimes completely unjustified. I've personally seen clusters of two medium machines running tasks that a laptop could have carried out even faster. Hence, always perform a descriptive/prescriptive analysis of the problem and the data, trying to focus on the following:

The current situation
Objectives (what do we need to achieve?)
Data and dimensionality (do we work with batch data? Do we have incoming streams?)
Acceptable delays (do we need real-time? Is it possible to process once a day/week?)

Big data solutions are justified, for example, when the following is the case:

The dataset cannot fit in the memory of a high-end machine
The incoming data flow is huge, continuous, and needs prompt computations (for example, clickstreams, web analytics, message dispatching, and so on)
It's not possible to split the data into small chunks because the acceptable delays are minimal (this piece of information must be mathematically quantified)
The operations can be parallelized efficiently (nowadays, many important algorithms have been implemented in distributed frameworks, but there are still tasks that cannot be processed by using parallel architectures)

In the chapter dedicated to recommendation systems, Chapter 12, Introduction to Recommendation Systems, we're going to discuss how to implement collaborative filtering using Apache Spark. The same framework will also be adopted for an example of Naive Bayes classification.

If you want to know more about the whole Hadoop ecosystem, visit http://hadoop.apache.org. Apache Mahout (http://mahout.apache.org) is a dedicated machine learning framework, and Spark (http://spark.apache.org), one the fastest computational engines, has a module called Machine Learning Library (MLlib) which implements many common algorithms that benefit from parallel processing.

Machine Learning Algorithms - Second Edition

Machine Learning Algorithms - Second Edition

Overview of this book

Related Content you might be interested in

Current Title:

Machine Learning Algorithms - Second Edition

Hands-On Unsupervised Learning with Python

Mastering Machine Learning Algorithms

Mastering Machine Learning Algorithms.