The Applied Data Science Workshop - Second Edition

By : Alex Galea

The Applied Data Science Workshop - Second Edition

By: Alex Galea

Overview of this book

From banking and manufacturing through to education and entertainment, using data science for business has revolutionized almost every sector in the modern world. It has an important role to play in everything from app development to network security. Taking an interactive approach to learning the fundamentals, this book is ideal for beginners. You’ll learn all the best practices and techniques for applying data science in the context of real-world scenarios and examples. Starting with an introduction to data science and machine learning, you’ll start by getting to grips with Jupyter functionality and features. You’ll use Python libraries like sci-kit learn, pandas, Matplotlib, and Seaborn to perform data analysis and data preprocessing on real-world datasets from within your own Jupyter environment. Progressing through the chapters, you’ll train classification models using sci-kit learn, and assess model performance using advanced validation techniques. Towards the end, you’ll use Jupyter Notebooks to document your research, build stakeholder reports, and even analyze web performance data. By the end of The Applied Data Science Workshop, you’ll be prepared to progress from being a beginner to taking your skills to the next level by confidently applying data science techniques and tools to real-world projects.

Preface

About the Book

Installing Libraries

1. Introduction to Jupyter Notebooks

Introduction

Basic Functionality and Features of Jupyter Notebooks

Jupyter Features

Summary

Free Chapter

2. Data Exploration with Jupyter

Introduction

Our First Analysis – the Boston Housing Dataset

Summary

3. Preparing Data for Predictive Modeling

Introduction

Machine Learning Process

Approaching Data Science Problems

Understanding Data from a Modeling Perspective

Introducing the Human Resource Analytics Dataset

Summary

4. Training Classification Models

Introduction

Understanding Classification Algorithms

Summary

5. Model Validation and Optimization

Introduction

Assessing Models with k-Fold Cross Validation

Dimensionality Reduction with PCA

Summary

6. Web Scraping with Jupyter Notebooks

Introduction

Internet Data Sources

Introduction to HTTP Requests

Data Workflow with pandas

Summary

Appendix

1. Introduction to Jupyter Notebooks

2. Data Exploration with Jupyter

3. Preparing Data for Predictive Modeling

4. Training Classification Models

5. Model Validation and Optimization

6. Web Scraping with Jupyter Notebooks

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Summary

In this chapter, we've gone over the basics of using Jupyter Notebooks for data science. We started by exploring the platform and finding our way around the interface. Then, we discussed the most useful features, which include tab completion and magic functions. Finally, we introduced the Python libraries we'll be using in this book.

As we'll see in the coming chapters, these libraries offer high-level abstractions that allow data science to be highly accessible with Python. This includes methods for creating statistical visualizations, building data cleaning pipelines, and training models on millions of data points and beyond.

While this chapter focused on the basics of Jupyter platforms, the next chapter is where the real data science begins. The remainder of this book is very interactive, and in Chapter 3, Preparing Data for Predictive Modeling, we'll perform an analysis of housing data using Jupyter Notebook and the Seaborn plotting library.

The Applied Data Science Workshop - Second Edition

By : Alex Galea

The Applied Data Science Workshop - Second Edition

By: Alex Galea

Overview of this book

Related Content you might be interested in

Current Title:

The Applied Data Science Workshop - Second Edition

scikit-learn Cookbook

The Machine Learning Workshop

The Supervised Learning Workshop

Summary