Learning Predictive Analytics with Python

Learning Predictive Analytics with Python

By : Ashish Kumar, Gary Dougan

Buy this Book

Learning Predictive Analytics with Python

By: Ashish Kumar, Gary Dougan

Buy this Book

Overview of this book

Social Media and the Internet of Things have resulted in an avalanche of data. Data is powerful but not in its raw form - It needs to be processed and modeled, and Python is one of the most robust tools out there to do so. It has an array of packages for predictive modeling and a suite of IDEs to choose from. Learning to predict who would win, lose, buy, lie, or die with Python is an indispensable skill set to have in this data age. This book is your guide to getting started with Predictive Analytics using Python. You will see how to process data and make predictive models from it. We balance both statistical and mathematical concepts, and implement them in Python using libraries such as pandas, scikit-learn, and numpy. You’ll start by getting an understanding of the basics of predictive modeling, then you will see how to cleanse your data of impurities and get it ready it for predictive modeling. You will also learn more about the best predictive modeling algorithms such as Linear Regression, Decision Trees, and Logistic Regression. Finally, you will see the best practices in predictive modeling, as well as the different applications of predictive modeling in the modern world.

Learning Predictive Analytics with Python

Credits

Foreword

About the Author

Acknowledgments

About the Reviewer

www.PacktPub.com

Preface

Free Chapter

Getting Started with Predictive Modelling

Introducing predictive modelling

Applications and examples of predictive modelling

Python and its packages – download and installation

Python and its packages for predictive modelling

IDEs for Python

Summary

Data Cleaning

Reading the data – variations and examples

Various methods of importing data in Python

The read_csv method

Use cases of the read_csv method

Case 2 – reading a dataset using the open method of Python

Case 3 – reading data from a URL

Case 4 – miscellaneous cases

Basics – summary, dimensions, and structure

Handling missing values

Creating dummy variables

Visualizing a dataset by basic plotting

Summary

Data Wrangling

Subsetting a dataset

Generating random numbers and their usage

Grouping the data – aggregation, filtering, and transformation

Random sampling – splitting a dataset in training and testing datasets

Concatenating and appending data

Merging/joining datasets

Summary

Statistical Concepts for Predictive Modelling

Random sampling and the central limit theorem

Hypothesis testing

Chi-square tests

Correlation

Summary

Linear Regression with Python

Understanding the maths behind linear regression

Making sense of result parameters

Implementing linear regression with Python

Model validation

Handling other issues in linear regression

Summary

Logistic Regression with Python

Linear regression versus logistic regression

Understanding the math behind logistic regression

Implementing logistic regression with Python

Model validation and evaluation

Model validation

Summary

Clustering with Python

Introduction to clustering – what, why, and how?

Mathematics behind clustering

Implementing clustering using Python

Fine-tuning the clustering

Summary

Trees and Random Forests with Python

Introducing decision trees

Understanding the mathematics behind decision trees

Implementing a decision tree with scikit-learn

Understanding and implementing regression trees

Understanding and implementing random forests

Summary

Best Practices for Predictive Modelling

Best practices for coding

Best practices for data handling

Best practices for algorithms

Best practices for statistics

Best practices for business contexts

Summary

A List of Links

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Model validation and evaluation

The preceding logistic regression model is built on the entire data. Let us now split the data into training and testing sets, build the model using the training set, and then check the accuracy using the testing set. The ultimate goal is to see whether it improves the accuracy of the prediction or not:

from sklearn.cross_validation import train_test_split
X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.3, random_state=0)

The preceding code snippet creates testing and training datasets for a predictor and also outcome variables. Let us now build a logistic regression model over the training set:

from sklearn import linear_model
from sklearn import metrics
clf1 = linear_model.LogisticRegression()
clf1.fit(X_train, Y_train)

The preceding code snippet creates the model. If you remember the equation behind the model, you will know that the model predicts probabilities and not the classes (binary output, that is, 0 or 1). One needs to select a...

Learning Predictive Analytics with Python

By : Ashish Kumar, Gary Dougan

Learning Predictive Analytics with Python

By: Ashish Kumar, Gary Dougan

Overview of this book

Related Content you might be interested in

Current Title:

Learning Predictive Analytics with Python

Model validation and evaluation