Statistical Application Development with R and Python

Statistical Application Development with R and Python - Second Edition

Overview of this book

Statistical Analysis involves collecting and examining data to describe the nature of data that needs to be analyzed. It helps you explore the relation of data and build models to make better decisions. This book explores statistical concepts along with R and Python, which are well integrated from the word go. Almost every concept has an R code going with it which exemplifies the strength of R and applications. The R code and programs have been further strengthened with equivalent Python programs. Thus, you will first understand the data characteristics, descriptive statistics and the exploratory attitude, which will give you firm footing of data analysis. Statistical inference will complete the technical footing of statistical methods. Regression, linear, logistic modeling, and CART, builds the essential toolkit. This will help you complete complex problems in the real world. You will begin with a brief understanding of the nature of data and end with modern and advanced statistical models like CART. Every step is taken with DATA and R code, and further enhanced by Python. The data analysis journey begins with exploratory analysis, which is more than simple, descriptive, data summaries. You will then apply linear regression modeling, and end with logistic regression, CART, and spatial statistics. By the end of this book you will be able to apply your statistical learning in major domains at work or in your projects.

Statistical Application Development with R and Python - Second Edition

Credits

About the Author

Acknowledgment

About the Reviewers

www.PacktPub.com

Customer Feedback

Preface

Free Chapter

Data Characteristics

Questionnaire and its components

Experiments with uncertainty in computer science

Installing and setting up R

Using R packages

Python installation and setup

IDEs for R and Python

The companion code bundle

Discrete distributions

Continuous distributions

Summary

Import/Export Data

Packages and settings – R and Python

Understanding data.frame and other formats

Using utils and the foreign packages

Exporting data/graphs

Pop quiz

Summary

Data Visualization

Packages and settings – R and Python

Visualization techniques for categorical data

Visualization techniques for continuous variable data

Pareto chart

A brief peek at ggplot2

Summary

Exploratory Analysis

Packages and settings – R and Python

Essential summary statistics

Techniques for exploratory analysis

Summary

Statistical Inference

Packages and settings – R and Python

Maximum likelihood estimator

Confidence intervals

Hypothesis testing

Summary

Linear Regression Analysis

Packages and settings - R and Python

The essence of regression

The simple linear regression model

Multiple linear regression model

Regression diagnostics

Model selection

Summary

Logistic Regression Model

Packages and settings – R and Python

Model validation and diagnostics

Logistic regression for the German credit screening dataset

Summary

Regression Models with Regularization

Packages and settings – R and Python

Regression spline

Ridge regression for linear models

Summary

Classification and Regression Trees

Packages and settings – R and Python

Splitting the data

Summary

CART and Beyond

Packages and settings – R and Python

Understanding bagging

Summary

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Regression diagnostics

In the Useful residual plots subsection, we saw how outliers can be identified using the residual plots. If there are outliers, we need to ask the following questions:

Is the observation an outlier due to an anomalous value in one or more covariate values?
Is the observation an outlier due to an extreme output value?
Is the observation an outlier because of both the covariate and output values being extreme values?

The distinction in the nature of an outlier is vital as one needs to be sure of its type. The techniques for an outlier identification are certainly different as is their impact. If the outlier is due to the covariate value, the observation is called a leverage point, and if it is due to the y value, we call it an influential point. The rest of the section is for the exact statistical technique for such an outlier identification.

Leverage points

As noted, a leverage point has an anomalous x value. The leverage points may be theoretically proved not to impact the...

Statistical Application Development with R and Python - Second Edition

Statistical Application Development with R and Python - Second Edition

Overview of this book

Related Content you might be interested in

Current Title:

Statistical Application Development with R and Python - Second Edition

Hands-On Ensemble Learning with R

Regression Analysis with R

Practical Data Science Cookbook, Second Edition

Regression diagnostics

Leverage points