Book Image

Reproducible Data Science with Pachyderm

By : Svetlana Karslioglu

Book Image

Reproducible Data Science with Pachyderm

By: Svetlana Karslioglu

Overview of this book

Pachyderm is an open source project that enables data scientists to run reproducible data pipelines and scale them to an enterprise level. This book will teach you how to implement Pachyderm to create collaborative data science workflows and reproduce your ML experiments at scale. You’ll begin your journey by exploring the importance of data reproducibility and comparing different data science platforms. Next, you’ll explore how Pachyderm fits into the picture and its significance, followed by learning how to install Pachyderm locally on your computer or a cloud platform of your choice. You’ll then discover the architectural components and Pachyderm's main pipeline principles and concepts. The book demonstrates how to use Pachyderm components to create your first data pipeline and advances to cover common operations involving data, such as uploading data to and from Pachyderm to create more complex pipelines. Based on what you've learned, you'll develop an end-to-end ML workflow, before trying out the hyperparameter tuning technique and the different supported Pachyderm language clients. Finally, you’ll learn how to use a SaaS version of Pachyderm with Pachyderm Notebooks. By the end of this book, you will learn all aspects of running your data pipelines in Pachyderm and manage them on a day-to-day basis.

Preface

Who this book is for

What this book covers

To get the most out of this book

Download the example code files

Download the color images

Conventions used

Share Your Thoughts

Section 1: Introduction to Pachyderm and Reproducible Data Science

Section 1: Introduction to Pachyderm and Reproducible Data Science

Free Chapter

Chapter 1: The Problem of Data Reproducibility

Chapter 1: The Problem of Data Reproducibility

Why is reproducibility important?

The reproducibility crisis in science

Demystifying MLOps

Types of data science platforms

Explaining ethical AI

Further reading

Chapter 2: Pachyderm Basics

Chapter 2: Pachyderm Basics

Reviewing Pachyderm architecture

Learning about version control primitives

Discovering pipeline elements

Further reading

Chapter 3: Pachyderm Pipeline Specification

Chapter 3: Pachyderm Pipeline Specification

Pipeline specification overview

Understanding inputs

Exploring informational parameters

Exploring transformation

Optimizing your pipeline

Exploring service parameters

Exploring output parameters

Further reading

Section 2:Getting Started with Pachyderm

Section 2:Getting Started with Pachyderm

Chapter 4: Installing Pachyderm Locally

Chapter 4: Installing Pachyderm Locally

Technical requirements

Installing the required tools

Installing minikube

Installing Docker Desktop

Installing the Pachyderm command-line interface

Enabling autocompletion for Pachyderm

Preparing the Kubernetes environment

Deploying Pachyderm

Accessing the Pachyderm Console

Deleting an existing Pachyderm deployment

Further reading

Chapter 5: Installing Pachyderm on a Cloud Platform

Chapter 5: Installing Pachyderm on a Cloud Platform

Technical requirements

Installing the required tools

Deploying Pachyderm on Amazon EKS

Deploying the cluster

Deploying Pachyderm on GKE

Deploying the cluster

Deploying Pachyderm on Microsoft AKS

Deploying the cluster

Accessing the Pachyderm console

Further reading

Chapter 6: Creating Your First Pipeline

Chapter 6: Creating Your First Pipeline

Technical requirements

Pipeline overview

Creating a repository

Creating a pipeline specification

Viewing the pipeline result

Adding another pipeline step

Further reading

Chapter 7: Pachyderm Operations

Chapter 7: Pachyderm Operations

Technical requirements

Reviewing the standard Pachyderm workflow

Executing data operations

Executing pipeline operations

Running maintenance operations

Further reading

Chapter 8: Creating an End-to-End Machine Learning Workflow

Chapter 8: Creating an End-to-End Machine Learning Workflow

Technical requirements

NLP example overview

Creating repositories and pipelines

Creating an NER pipeline

Retraining an NER model

Further reading

Chapter 9: Distributed Hyperparameter Tuning with Pachyderm

Chapter 9: Distributed Hyperparameter Tuning with Pachyderm

Technical requirements

Reviewing hyperparameter tuning techniques and strategies

Creating a hyperparameter tuning pipeline in Pachyderm

Further reading

Section 3:Pachyderm Clients and Tools

Section 3:Pachyderm Clients and Tools

Chapter 10: Pachyderm Language Clients

Chapter 10: Pachyderm Language Clients

Technical requirements

Using the Pachyderm Go client

Using the Pachyderm Python client

Further reading

Chapter 11: Using Pachyderm Notebooks

Chapter 11: Using Pachyderm Notebooks

Technical requirements

Enabling Pachyderm Notebooks in Pachyderm Hub

Running basic Pachyderm operations in Pachyderm Notebooks

Creating and running an example pipeline in Pachyderm Notebooks

Further reading

Other Books You May Enjoy

Other Books You May Enjoy

Packt is searching for authors like you

Share Your Thoughts

Customer Reviews

5 star

0

4 star

0

3 star

0

2 star

0

1 star

0

Summary

In this chapter, we have discussed a number of important concepts that help define why reproducibility is important and why it should be a part of a successful data science process.

We've learned that data science models are used to analyze historical data as input with a target goal to calculate the most probable and most successful result. We've established that replication, the ability to reproduce the results of a scientific experiment, is one of the fundamental principles of good research and that it is one of the best ways to ensure that your team is doing everything to reduce bias in your models. Bias can creep into a calculation from misrepresentation in a training dataset. Often, this reflects historical and social realities and norms accepted in society. Another way to reduce bias in your training data is to have a diverse team that includes representatives of all genders, races, and backgrounds.

We've learned that data dredging, or fishing, is an unethical technique used by some data scientists to prove a predefined hypothesis by cherry-picking the results of an experiment and only selecting the results that prove the desired outcome and ignoring any inconvenient trends.

We've also learned about the MLOps methodology, a lifecycle of a machine learning application, similar in its principle to the DevOps software lifecycle technique. MLOps includes the following main phases: planning, development, training, validation, deployment, and monitoring. All of the phases are continuously repeated, creating a feedback loop that ensures seamless experiment management from planning through development and testing to production and post-production phases.

We've also reviewed some of the most important aspects of ethical AI, a discipline of data science that focuses on ethical aspects of artificial intelligence, robotics, and data science. A failure to implement an ethical AI process in your organization might lead to undesirable legal consequences if deployed production models are found to be discriminatory.

In the next chapter, we will learn about the main concepts of the Pachyderm version-control system, which can help you address many of the issues described in this chapter.