Hands-On Data Science with Anaconda

By : Yuxing Yan, James Yan

Hands-On Data Science with Anaconda

By: Yuxing Yan, James Yan

Overview of this book

Anaconda is an open source platform that brings together the best tools for data science professionals with more than 100 popular packages supporting Python, Scala, and R languages. Hands-On Data Science with Anaconda gets you started with Anaconda and demonstrates how you can use it to perform data science operations in the real world. The book begins with setting up the environment for Anaconda platform in order to make it accessible for tools and frameworks such as Jupyter, pandas, matplotlib, Python, R, Julia, and more. You’ll walk through package manager Conda, through which you can automatically manage all packages including cross-language dependencies, and work across Linux, macOS, and Windows. You’ll explore all the essentials of data science and linear algebra to perform data science tasks using packages such as SciPy, contrastive, scikit-learn, Rattle, and Rmixmod. Once you’re accustomed to all this, you’ll start with operations in data science such as cleaning, sorting, and data classification. You’ll move on to learning how to perform tasks such as clustering, regression, prediction, and building machine learning models and optimizing them. In addition to this, you’ll learn how to visualize data using the packages available for Julia, Python, and R.

Preface

Who this book is for

What this book covers

To get the most out of this book

Get in touch

Free Chapter

Ecosystem of Anaconda

Summary

Review questions and exercises

Anaconda Installation

Installing Anaconda

Testing Python

Using IPython

Using Python via Jupyter

Introducing Spyder

Installing R via Conda

Installing Julia and linking it to Jupyter

Installing Octave and linking it to Jupyter

Finding help

Summary

Review questions and exercises

Data Basics

Sources of data

UCI machine learning

Introduction to the Python pandas package

Several ways to input data

Introduction to the Quandl data delivery platform

Dealing with missing data

Data sorting

Introduction to the cbsodata Python package

Introduction to the datadotworld Python package

Introduction to the haven and foreign R packages

Introduction to the dslabs R package

Generating Python datasets

Generating R datasets

Summary

Review questions and exercises

Data Visualization

Importance of data visualization

Data visualization in R

Data visualization in Python

Data visualization in Julia

Drawing simple graphs

Visualization packages for R

Visualization packages for Python

Visualization packages for Julia

Dynamic visualization

Summary

Review questions and exercises

Statistical Modeling in Anaconda

Introduction to linear models

Running a linear regression in R, Python, Julia, and Octave

Critical value and the decision rule

F-test, critical value, and the decision rule

Dealing with missing data

Detecting outliers and treatments

Several multivariate linear models

Collinearity and its solution

A model's performance measure

Summary

Review questions and exercises

Managing Packages

Introduction to packages, modules, or toolboxes

Two examples of using packages

Finding all R packages

Finding all Python packages

Finding all Julia packages

Finding all Octave packages

Task views for R

Finding manuals

Package dependencies

Package management in R

Package management in Python

Package management in Julia

Package management in Octave

Conda – the package manager

Creating a set of programs in R and Python

Finding environmental variables

Summary

Review questions and exercises

Optimization in Anaconda

Why optimization is important

General issues for optimization problems

Quadratic optimization

Example #1 – stock portfolio optimization

Example #2 – optimal tax policy

Packages for optimization in R

Packages for optimization in Python

Packages for optimization in Octave

Packages for optimization in Julia

Summary

Review questions and exercises

Unsupervised Learning in Anaconda

Introduction to unsupervised learning

Hierarchical clustering

k-means clustering

Introduction to Python packages – scipy

Introduction to Python packages – contrastive

Introduction to Python packages – sklearn (scikit-learn)

Introduction to R packages – rattle

Introduction to R packages – randomUniformForest

Introduction to R packages – Rmixmod

Implementation using Julia

Task view for Cluster Analysis

Summary

Review questions and exercises

Supervised Learning in Anaconda

A glance at supervised learning

Classification

Implementation of supervised learning via R

Implementation via Python

Implementation via Octave

Implementation via Julia

Summary

Review questions and exercises

Predictive Data Analytics – Modeling and Validation

Understanding predictive data analytics

Useful datasets

Predicting future events

Model selection

Granger causality test

Summary

Review questions and exercises

Anaconda Cloud

Introduction to Anaconda Cloud

Jupyter Notebook in depth

Replicating others' environments locally

Summary

Review questions and exercises

Distributed Computing, Parallel Computing, and HPCC

Introduction to distributed versus parallel computing

Understanding MPI

Parallel processing in Python

Compute nodes

Anaconda add-on

Introduction to HPCC

Summary

Review questions and exercises

References

Chapter 01: Ecosystem of Anaconda

Chapter 02: Anaconda Installation

Chapter 03: Data Basics

Chapter 04: Data Visualization

Chapter 05: Statistical Modeling in Anaconda

Chapter 06: Managing Packages

Chapter 07: Optimization in Anaconda

Chapter 08: Unsupervised Learning in Anaconda

Chapter 09: Supervised Learning in Anaconda

Chapter 10: Predictive Data Analytics – Modelling and Validation

Chapter 11: Anaconda Cloud

Chapter 12: Distributed Computing, Parallel Computing, and HPCC

Other Books You May Enjoy

Leave a review - let other readers know what you think

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Preface

Anaconda is an open source data science platform that brings the best tools for data science together. It is a data science stack that includes more than 100 popular packages based on Python, Scala, and R. With the help of its package manager, conda, users can work with hundreds of packages in different languages and perform data preprocessing, modeling, clustering, classification, and validation with ease.

This book will get you started with Anaconda and how you can use it to perform data science operations in the real world. You will start of setting up the environment for the Anaconda platform, Jupyter, and installing the relevant packages. You will then cover the basics of data science and linear algebra for performing data science tasks. Once you are ready to go, you will start with data science operations such as cleaning, sorting, and data classification. You will then learn how to perform tasks such as clustering, regression, prediction, building machine learning models, and optimizing them. You will also learn how to visualize data and share the projects.

During this course, you will learn how to use different packages, using Anaconda to get the best results. You will learn how to efficiently use conda — the package, dependency, and environment manager for Anaconda. You will also be introduced to several powerful features of Anaconda, such as additional projects, project add-ons, shared project drives, and powerful compute nodes that are available in the paid version for accomplishing advanced data handling processes. You will learn how to build scalable and functionally efficient packages, and how to perform heterogeneous data exploration, distributed computing, and more. You will learn to discover and share packages, notebooks, and environments to increase productivity. You will also learn about Anaconda Accelerate, a feature that can help you to achieve SLAs easily and optimize computational power.

In this book, we introduce four programming languages: R, Python, Octave, and Julia. There are several reasons for doing so. Firstly, all four are open source, which is one of the future trends. Secondly, one of the most obvious advantages to using the Anaconda platform is that it allows you to where we could implement many programs written in different languages. However, for many new readers, learning four languages at the same time would be quite challenging. The best strategy is to focus on R and Python first. After a while, or after finishing the whole book, learn Octave or Julia on the second reading.

R: This is a free software environment for statistical computing and graphics. It compiles and runs on a wide variety of UNIX platforms, such as Windows and macOS. We think that R might be the easiest of many good computer languages, especially those that offer free software. The author has published a book entitled Financial Modeling using R; you can refer to its Amazon link at http://canisius.edu/~yany/webs/amazon2018R.shtml.
Python: This is an interpreted high-level programming language for general-purpose programming. For business analytics/data science, Python is probably the number 1 choice out of many promising computer languages. In 2017, the author published a book entitled Python for Finance (second edition); you can refer to its Amazon link at http://canisius.edu/~yany/webs/amazonP4F2.shtml.
Octave: This is a piece of software featuring a high-level programming language, primarily intended for numerical computations. Octave helps with solving linear and nonlinear problems numerically, as well as performing other numerical experiments. Octave is also free. Its syntax is largely compatible with MATLAB, which is quite popular on Wall Street and in other industries.
Julia: This is a high-level, high-performance dynamic programming language for numerical computing. It provides a sophisticated compiler, distributed parallel execution, numerical accuracy, and an extensive mathematical function library. Julia’s base library, largely written in Julia itself, also integrates mature, best-of-breed, open source C and Fortran libraries for linear algebra, random number generation, signal processing, and string processing.

Happy reading!

Hands-On Data Science with Anaconda

By : Yuxing Yan, James Yan

Hands-On Data Science with Anaconda

By: Yuxing Yan, James Yan

Overview of this book

Related Content you might be interested in

Current Title:

Hands-On Data Science with Anaconda

Python for Finance

Learning Quantitative Finance with R

Python for Finance Cookbook