Applied Unsupervised Learning with Python

Applied Unsupervised Learning with Python

By : Benjamin Johnston, Aaron Jones, Christopher Kruger

Buy this Book

Applied Unsupervised Learning with Python

By: Benjamin Johnston, Aaron Jones, Christopher Kruger

Buy this Book

Overview of this book

Unsupervised learning is a useful and practical solution in situations where labeled data is not available. Applied Unsupervised Learning with Python guides you in learning the best practices for using unsupervised learning techniques in tandem with Python libraries and extracting meaningful information from unstructured data. The book begins by explaining how basic clustering works to find similar data points in a set. Once you are well-versed with the k-means algorithm and how it operates, you’ll learn what dimensionality reduction is and where to apply it. As you progress, you’ll learn various neural network techniques and how they can improve your model. While studying the applications of unsupervised learning, you will also understand how to mine topics that are trending on Twitter and Facebook and build a news recommendation engine for users. Finally, you will be able to put your knowledge to work through interesting activities such as performing a Market Basket Analysis and identifying relationships between different products. By the end of this book, you will have the skills you need to confidently build your own models using Python.

Applied Unsupervised Learning with Python

Preface

Free Chapter

Introduction to Clustering

Introduction

Unsupervised Learning versus Supervised Learning

Clustering

Introduction to k-means Clustering

Activity 1: Implementing k-means Clustering

Summary

Hierarchical Clustering

Introduction

Clustering Refresher

The Organization of Hierarchy

Introduction to Hierarchical Clustering

Linkage

Agglomerative versus Divisive Clustering

k-means versus Hierarchical Clustering

Summary

Neighborhood Approaches and DBSCAN

Introduction

Introduction to DBSCAN

DBSCAN Versus k-means and Hierarchical Clustering

Summary

Dimension Reduction and PCA

Introduction

Overview of Dimensionality Reduction Techniques

PCA

Summary

Autoencoders

Introduction

Fundamentals of Artificial Neural Networks

Autoencoders

Summary

t-Distributed Stochastic Neighbor Embedding (t-SNE)

Introduction

Stochastic Neighbor Embedding (SNE)

t-Distributed SNE

Interpreting t-SNE Plots

Summary

Topic Modeling

Introduction

Cleaning Text Data

Latent Dirichlet Allocation

Non-Negative Matrix Factorization

Summary

Market Basket Analysis

Introduction

Market Basket Analysis

Characteristics of Transaction Data

Apriori Algorithm

Association Rules

Summary

Hotspot Analysis

Introduction

Kernel Density Estimation

Hotspot Analysis

Summary

Appendix

Chapter 1: Introduction to Clustering

Chapter 2: Hierarchical Clustering

Chapter 3: Neighborhood Approaches and DBSCAN

Chapter 4: Dimension Reduction and PCA

Chapter 5: Autoencoders

Chapter 6: t-Distributed Stochastic Neighbor Embedding (t-SNE)

Chapter 7: Topic Modeling

Chapter 8: Market Basket Analysis

Chapter 9: Hotspot Analysis

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Unsupervised Learning versus Supervised Learning

Unsupervised learning is one of the most exciting areas of development in machine learning today. If you have explored machine learning bookwork before, you are probably familiar with the common breakout of problems in either supervised or unsupervised learning. Supervised learning encompasses the problem set of having a labeled dataset that can be used to either classify (for example, predicting smokers and non-smokers if you're looking at a lung health dataset) or fit a regression line on (for example, predicting the sale price of a home based on how many bedrooms it has). This model most closely mirrors an intuitive human approach to learning.

If you wanted to learn how to not burn your food with a basic understanding of cooking, you could build a dataset by putting your food on the burner and seeing how long it takes (input) for your food to burn (output). Eventually, as you continue to burn your food, you will build a mental model of when burning will occur and avoid it in the future. Development in supervised learning was once fast-paced and valuable, but it has since simmered down in recent years – many of the obstacles to knowing your data have already been tackled:

Figure 1.1: Differences between unsupervised and supervised learning

Conversely, unsupervised learning encompasses the problem set of having a tremendous amount of data that is unlabeled. Labeled data, in this case, would be data that has a supplied "target" outcome that you are trying to find the correlation to with supplied data (you know that you are looking for whether your food was burned in the preceding example). Unlabeled data is when you do not know what the "target" outcome is, and you only have supplied input data.

Building upon the previous example, imagine you were just dropped on planet Earth with zero knowledge of how cooking works. You are given 100 days, a stove, and a fridge full of food without any instructions on what to do. Your initial exploration of a kitchen could go in infinite directions – on day 10, you may finally learn how to open the fridge; on day 30, you may learn that food can go on the stove; and after many more days, you may unwittingly make an edible meal. As you can see, trying to find meaning in a kitchen devoid of adequate informational structure leads to very noisy data that is completely irrelevant to actually preparing a meal.

Unsupervised learning can be an answer to this problem. By looking back at your 100 days of data, clustering can be used to find patterns of similar days where a meal was produced, and you can easily review what you did on those days. However, unsupervised learning isn't a magical answer –simply finding clusters can be just as likely to help you to find pockets of similar yet ultimately useless data.

This challenge is what makes unsupervised learning so exciting. How can we find smarter techniques to speed up the process of finding clusters of information that are beneficial to our end goals?

Applied Unsupervised Learning with Python

By : Benjamin Johnston, Aaron Jones, Christopher Kruger

Applied Unsupervised Learning with Python

By: Benjamin Johnston, Aaron Jones, Christopher Kruger

Overview of this book

Related Content you might be interested in

Current Title:

Applied Unsupervised Learning with Python

Applied Unsupervised Learning with R

Hands-On Unsupervised Learning with Python

Data Science with Python

Unsupervised Learning versus Supervised Learning