Principles of Data Science - Second Edition

By : Sinan Ozdemir, Sunil Kakade, Marco Tibaldeschi

Principles of Data Science - Second Edition

By: Sinan Ozdemir, Sunil Kakade, Marco Tibaldeschi

Overview of this book

Need to turn programming skills into effective data science skills? This book helps you connect mathematics, programming, and business analysis. You’ll feel confident asking—and answering—complex, sophisticated questions of your data, making abstract and raw statistics into actionable ideas. Going through the data science pipeline, you'll clean and prepare data and learn effective data mining strategies and techniques to gain a comprehensive view of how the data science puzzle fits together. You’ll learn fundamentals of computational mathematics and statistics and pseudo-code used by data scientists and analysts. You’ll learn machine learning, discovering statistical models that help control and navigate even the densest datasets, and learn powerful visualizations that communicate what your data means.

Preface

Who this book is for

What this book covers

To get the most out of this book

Get in touch

Free Chapter

1. How to Sound Like a Data Scientist

What is data science?

The data science Venn diagram

Why Python?

Some more terminology

Data science case studies

Summary

2. Types of Data

Flavors of data

Why look at these distinctions?

Structured versus unstructured data

Quantitative versus qualitative data

The road thus far

The four levels of data

Data is in the eye of the beholder

Summary

3. The Five Steps of Data Science

Introduction to data science

Overview of the five steps

Exploring the data

Summary

4. Basic Mathematics

Mathematics as a discipline

Basic symbols and terminology

Linear algebra

Summary

5. Impossible or Improbable - A Gentle Introduction to Probability

Basic definitions

Probability

Bayesian versus Frequentist

Compound events

Conditional probability

The rules of probability

A bit deeper

Summary

6. Advanced Probability

Collectively exhaustive events

Bayesian ideas revisited

Random variables

Summary

7. Basic Statistics

What are statistics?

How do we obtain and sample data?

8. Advanced Statistics

Point estimates

Sampling distributions

Confidence intervals

Hypothesis tests

Summary

9. Communicating Data

Why does communication matter?

Identifying effective and ineffective visualizations

When graphs and statistics lie

Verbal communication

The why/how/what strategy of presenting

Summary

10. How to Tell If Your Toaster Is Learning – Machine Learning Essentials

What is machine learning?

Machine learning isn't perfect

How does machine learning work?

Types of machine learning

How does statistical modeling fit into all of this?

Linear regression

Logistic regression

Probability, odds, and log odds

Dummy variables

Summary

11. Predictions Don't Grow on Trees - or Do They?

Naive Bayes classification

Decision trees

Unsupervised learning

k-means clustering

Choosing an optimal number for K and cluster validation

Summary

12. Beyond the Essentials

The bias/variance trade-off

K folds cross-validation

Grid searching

Ensembling techniques

Neural networks

Summary

13. Case Studies

Case study 1 – Predicting stock prices based on social media

Case study 2 – Why do some people cheat on their spouses?

Case study 3 – Using TensorFlow

Summary

14. Building Machine Learning Models with Azure Databricks and Azure Machine Learning service

Technical requirements

Technologies for machine learning projects

Configuring Azure Databricks

Training a text classifier with Azure Databricks

Azure Machine Learning

Summary

Other Books You May Enjoy

Leave a review – let other readers know what you think

Index

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Chapter 1. How to Sound Like a Data Scientist

No matter which industry you work in—IT, fashion, food, or finance—there is no doubt that data affects your life and work. At some point this week, you will either have or hear a conversation about data. News outlets are covering more and more stories about data leaks, cybercrimes, and how data can give us a glimpse into our lives. But why now? What makes this era such a hotbed of data-related industries?

In the nineteenth century, the world was in the grip of the Industrial Age. Mankind was exploring its place in the industrial world, working with giant mechanical inventions. Captains of industry, such as Henry Ford, recognized that using these machines could open major market opportunities, enabling industries to achieve previously unimaginable profits. Of course, the Industrial Age had its pros and cons. While mass production placed goods in the hands of more consumers, our battle with pollution also began at around this time.

By the twentieth century, we were quite skilled at making huge machines; the goal now was to make them smaller and faster. The Industrial Age was over and was replaced by what we now refer to as the Information Age. We started using machines to gather and store information (data) about ourselves and our environment for the purpose of understanding our universe.

Beginning in the 1940s, machines such as ENIAC (considered one of the first—if not the first—computers) were computing math equations and running models and simulations like never before. The following photograph shows ENIAC:

ENIAC—The world's first electronic digital computer (Ref: http://ftp.arl.mil/ftp/historic-computers/)

We finally had a decent lab assistant who could run the numbers better than we could! As with the Industrial Age, the Information Age brought us both the good and the bad. The good was the extraordinary works of technology, including mobile phones and televisions. The bad was not as bad as worldwide pollution, but still left us with a problem in the twenty-first century—so much data.

That's right—the Information Age, in its quest to procure data, has exploded the production of electronic data. Estimates show that we created about 1.8 trillion gigabytes of data in 2011 (take a moment to just think about how much that is). Just one year later, in 2012, we created over 2.8 trillion gigabytes of data! This number is only going to explode further to hit an estimated 40 trillion gigabytes of created data in just one year by 2020. People contribute to this every time they tweet, post on Facebook, save a new resume on Microsoft Word, or just send their mom a picture by text message.

Not only are we creating data at an unprecedented rate, but we are also consuming it at an accelerated pace as well. Just five years ago, in 2013, the average cell phone user used under 1 GB of data a month. Today, that number is estimated to be well over 2 GB a month. We aren't just looking for the next personality quiz—what we are looking for is insight. With all of this data out there, some of it has to be useful to me! And it can be!

So we, in the twenty-first century, are left with a problem. We have so much data and we keep making more. We have built insanely tiny machines that collect data 24/7, and it's our job to make sense of it all. Enter the Data Age. This is the age when we take machines dreamed up by our nineteenth century ancestors and the data created by our twentieth century counterparts and create insights and sources of knowledge that every human on Earth can benefit from. The United States created an entirely new role in the government of chief data scientist. Many companies are now investing in data science departments and hiring data scientists. The benefit is quite obvious—using data to make accurate predictions and simulations gives us insight into our world like never before.

Sounds great, but what's the catch?

This chapter will explore the terminology and vocabulary of the modern data scientist. We will learn keywords and phrases that will be essential in our discussion of data science throughout this book. We will also learn why we use data science and learn about the three key domains that data science is derived from before we begin to look at the code in Python, the primary language used in this book. This chapter will cover the following topics:

The basic terminology of data science
The three domains of data science
The basic Python syntax

Principles of Data Science - Second Edition

By : Sinan Ozdemir, Sunil Kakade, Marco Tibaldeschi

Principles of Data Science - Second Edition

By: Sinan Ozdemir, Sunil Kakade, Marco Tibaldeschi

Overview of this book

Related Content you might be interested in

Current Title:

Principles of Data Science - Second Edition

Chapter 1. How to Sound Like a Data Scientist