Mastering Predictive Analytics with R

Mastering Predictive Analytics with R - Second Edition

By : James D. Miller, Rui Miguel Forte

Buy this Book

Mastering Predictive Analytics with R - Second Edition

By: James D. Miller, Rui Miguel Forte

Buy this Book

Overview of this book

R offers a free and open source environment that is perfect for both learning and deploying predictive modeling solutions. With its constantly growing community and plethora of packages, R offers the functionality to deal with a truly vast array of problems. The book begins with a dedicated chapter on the language of models and the predictive modeling process. You will understand the learning curve and the process of tidying data. Each subsequent chapter tackles a particular type of model, such as neural networks, and focuses on the three important questions of how the model works, how to use R to train it, and how to measure and assess its performance using real-world datasets. How do you train models that can handle really large datasets? This book will also show you just that. Finally, you will tackle the really important topic of deep learning by implementing applications on word embedding and recurrent neural networks. By the end of this book, you will have explored and tested the most popular modeling techniques in use on real- world datasets and mastered a diverse range of techniques in predictive analytics using R.

Mastering Predictive Analytics with R Second Edition

Credits

About the Authors

About the Reviewer

www.PacktPub.com

Customer Feedback

Preface

Free Chapter

Gearing Up for Predictive Modeling

Models

Types of model

The process of predictive modeling

Summary

Tidying Data and Measuring Performance

Getting started

Tidying data

Categorizing data quality

Performance metrics

Cross-validation

Learning curves

Summary

Linear Regression

Introduction to linear regression

Simple linear regression

Multiple linear regression

Assessing linear regression models

Problems with linear regression

Feature selection

Regularization

Polynomial regression

Summary

Generalized Linear Models

Classifying with linear regression

Introduction to logistic regression

Predicting heart disease

Assessing logistic regression models

Regularization with the lasso

Classification metrics

Extensions of the binary logistic classifier

Poisson regression

Negative Binomial regression

Summary

Neural Networks

The biological neuron

The artificial neuron

Stochastic gradient descent

Multilayer perceptron networks

The back propagation algorithm

Predicting the energy efficiency of buildings

Predicting glass type revisited

Predicting handwritten digits

Radial basis function networks

Summary

Support Vector Machines

Maximal margin classification

Support vector classification

Kernels and support vector machines

Predicting chemical biodegration

Predicting credit scores

Multiclass classification with support vector machines

Summary

Tree-Based Methods

The intuition for tree models

Algorithms for training decision trees

Predicting class membership on synthetic 2D data

Predicting the authenticity of banknotes

Predicting complex skill learning

Improvements to the M5 model

Summary

Dimensionality Reduction

Defining DR

Summary

Ensemble Methods

Bagging

Boosting

Predicting atmospheric gamma ray radiation

Predicting complex skill learning with boosting

Summary

Probabilistic Graphical Models

A little graph theory

Bayes' theorem

Conditional independence

Bayesian networks

The Naïve Bayes classifier

Summary

Topic Modeling

An overview of topic modeling

Latent Dirichlet Allocation

Modeling the topics of online news stories

Modeling tweet topics

Summary

Recommendation Systems

Rating matrix

Collaborative filtering

Singular value decomposition

Predicting recommendations for movies and jokes

Loading and pre-processing the data

Exploring the data

Other approaches to recommendation systems

Summary

Scaling Up

Starting the project

Characteristics of big data

Training models at scale

A path forward

Alternatives

Summary

Deep Learning

Machine learning or deep learning

What is deep learning?

Summary

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Predicting complex skill learning with boosting

We will revisit our Skillcraft dataset in this section--this time in the context of another boosting technique known as stochastic gradient boosting. The main characteristic of this method is that in every iteration of boosting, we compute a gradient in the direction of the errors that are made by the model trained in the current iteration.

This gradient is then used in order to guide the construction of the model that will be added in the next iteration. Stochastic gradient boosting is commonly used with decision trees, and a good implementation in R can be found in the gbm package, which provides us with the gbm() function. For regression problems, we need to specify the distribution parameter to be gaussian. In addition, we can specify the number of trees we want to build (which is equivalent to the number of iterations of boosting) via the n.trees parameter, as well as a shrinkage parameter that is used to control the algorithm's learning...

Mastering Predictive Analytics with R - Second Edition

By : James D. Miller, Rui Miguel Forte

Mastering Predictive Analytics with R - Second Edition

By: James D. Miller, Rui Miguel Forte

Overview of this book

Related Content you might be interested in

Current Title:

Mastering Predictive Analytics with R - Second Edition

Statistics for Data Science

Mastering Machine Learning with R

Mastering Machine Learning with R.

Predicting complex skill learning with boosting