The Unsupervised Learning Workshop

By : Aaron Jones, Christopher Kruger, Benjamin Johnston

The Unsupervised Learning Workshop

By: Aaron Jones, Christopher Kruger, Benjamin Johnston

Overview of this book

Do you find it difficult to understand how popular companies like WhatsApp and Amazon find valuable insights from large amounts of unorganized data? The Unsupervised Learning Workshop will give you the confidence to deal with cluttered and unlabeled datasets, using unsupervised algorithms in an easy and interactive manner. The book starts by introducing the most popular clustering algorithms of unsupervised learning. You'll find out how hierarchical clustering differs from k-means, along with understanding how to apply DBSCAN to highly complex and noisy data. Moving ahead, you'll use autoencoders for efficient data encoding. As you progress, you’ll use t-SNE models to extract high-dimensional information into a lower dimension for better visualization, in addition to working with topic modeling for implementing natural language processing (NLP). In later chapters, you’ll find key relationships between customers and businesses using Market Basket Analysis, before going on to use Hotspot Analysis for estimating the population density of an area. By the end of this book, you’ll be equipped with the skills you need to apply unsupervised algorithms on cluttered datasets to find useful patterns and insights.

Preface

About the Book

1. Introduction to Clustering

Introduction

Unsupervised Learning versus Supervised Learning

Clustering

Introduction to k-means Clustering

Summary

Free Chapter

2. Hierarchical Clustering

Introduction

Clustering Refresher

The Organization of the Hierarchy

Introduction to Hierarchical Clustering

Linkage

Agglomerative versus Divisive Clustering

k-means versus Hierarchical Clustering

Summary

3. Neighborhood Approaches and DBSCAN

Introduction

Clusters as Neighborhoods

Introduction to DBSCAN

DBSCAN versus k-means and Hierarchical Clustering

Summary

4. Dimensionality Reduction Techniques and PCA

Introduction

What Is Dimensionality Reduction?

Overview of Dimensionality Reduction Techniques

Principal Component Analysis

Summary

5. Autoencoders

Introduction

Fundamentals of Artificial Neural Networks

Autoencoders

Summary

6. t-Distributed Stochastic Neighbor Embedding

Introduction

The MNIST Dataset

Stochastic Neighbor Embedding (SNE)

t-Distributed SNE

Interpreting t-SNE Plots

Summary

7. Topic Modeling

Introduction

Topic Models

Cleaning Text Data

Latent Dirichlet Allocation

Non-Negative Matrix Factorization

Summary

8. Market Basket Analysis

Introduction

Market Basket Analysis

Characteristics of Transaction Data

The Apriori Algorithm

Association Rules

Summary

9. Hotspot Analysis

Introduction

Spatial Statistics

Kernel Density Estimation

Hotspot Analysis

Summary

Appendix

1. Introduction to Clustering

2. Hierarchical Clustering

3. Neighborhood Approaches and DBSCAN

4. Dimensionality Reduction Techniques and PCA

5. Autoencoders

6. t-Distributed Stochastic Neighbor Embedding

7. Topic Modeling

8. Market Basket Analysis

9. Hotspot Analysis

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Stochastic Neighbor Embedding (SNE)

SNE is one of a number of different methods that fall within the category of manifold learning, which aims to describe high-dimensional spaces within low-dimensional manifolds or bounded areas. At first thought, this seems like an impossible task; how can we reasonably represent data in two dimensions if we have a dataset with at least 30 features? As we work through the derivation of SNE, it is hoped that you will see how this is possible. Don't worry – we will not be covering the mathematical details of this process in great depth as it is outside of the scope of this chapter. Constructing an SNE can be divided into the following steps:

Convert the distances between datapoints in the high-dimensional space into conditional probabilities. Say we had two points, xi and xj, in a high-dimensional space and we wanted to determine the probability (pi|j) that xj would be picked as a neighbor of xi. To define this probability, we use...

The Unsupervised Learning Workshop

By : Aaron Jones, Christopher Kruger, Benjamin Johnston

The Unsupervised Learning Workshop

By: Aaron Jones, Christopher Kruger, Benjamin Johnston

Overview of this book

Related Content you might be interested in

Current Title:

The Unsupervised Learning Workshop

Applied Unsupervised Learning with R

Hands-On Unsupervised Learning with Python

The Machine Learning Workshop

Stochastic Neighbor Embedding (SNE)