Book Image

Reproducible Data Science with Pachyderm

By : Svetlana Karslioglu

Book Image

Reproducible Data Science with Pachyderm

By: Svetlana Karslioglu

Overview of this book

Pachyderm is an open source project that enables data scientists to run reproducible data pipelines and scale them to an enterprise level. This book will teach you how to implement Pachyderm to create collaborative data science workflows and reproduce your ML experiments at scale. You’ll begin your journey by exploring the importance of data reproducibility and comparing different data science platforms. Next, you’ll explore how Pachyderm fits into the picture and its significance, followed by learning how to install Pachyderm locally on your computer or a cloud platform of your choice. You’ll then discover the architectural components and Pachyderm's main pipeline principles and concepts. The book demonstrates how to use Pachyderm components to create your first data pipeline and advances to cover common operations involving data, such as uploading data to and from Pachyderm to create more complex pipelines. Based on what you've learned, you'll develop an end-to-end ML workflow, before trying out the hyperparameter tuning technique and the different supported Pachyderm language clients. Finally, you’ll learn how to use a SaaS version of Pachyderm with Pachyderm Notebooks. By the end of this book, you will learn all aspects of running your data pipelines in Pachyderm and manage them on a day-to-day basis.

Preface

Who this book is for

What this book covers

To get the most out of this book

Download the example code files

Download the color images

Conventions used

Share Your Thoughts

Section 1: Introduction to Pachyderm and Reproducible Data Science

Section 1: Introduction to Pachyderm and Reproducible Data Science

Free Chapter

Chapter 1: The Problem of Data Reproducibility

Chapter 1: The Problem of Data Reproducibility

Why is reproducibility important?

The reproducibility crisis in science

Demystifying MLOps

Types of data science platforms

Explaining ethical AI

Further reading

Chapter 2: Pachyderm Basics

Chapter 2: Pachyderm Basics

Reviewing Pachyderm architecture

Learning about version control primitives

Discovering pipeline elements

Further reading

Chapter 3: Pachyderm Pipeline Specification

Chapter 3: Pachyderm Pipeline Specification

Pipeline specification overview

Understanding inputs

Exploring informational parameters

Exploring transformation

Optimizing your pipeline

Exploring service parameters

Exploring output parameters

Further reading

Section 2:Getting Started with Pachyderm

Section 2:Getting Started with Pachyderm

Chapter 4: Installing Pachyderm Locally

Chapter 4: Installing Pachyderm Locally

Technical requirements

Installing the required tools

Installing minikube

Installing Docker Desktop

Installing the Pachyderm command-line interface

Enabling autocompletion for Pachyderm

Preparing the Kubernetes environment

Deploying Pachyderm

Accessing the Pachyderm Console

Deleting an existing Pachyderm deployment

Further reading

Chapter 5: Installing Pachyderm on a Cloud Platform

Chapter 5: Installing Pachyderm on a Cloud Platform

Technical requirements

Installing the required tools

Deploying Pachyderm on Amazon EKS

Deploying the cluster

Deploying Pachyderm on GKE

Deploying the cluster

Deploying Pachyderm on Microsoft AKS

Deploying the cluster

Accessing the Pachyderm console

Further reading

Chapter 6: Creating Your First Pipeline

Chapter 6: Creating Your First Pipeline

Technical requirements

Pipeline overview

Creating a repository

Creating a pipeline specification

Viewing the pipeline result

Adding another pipeline step

Further reading

Chapter 7: Pachyderm Operations

Chapter 7: Pachyderm Operations

Technical requirements

Reviewing the standard Pachyderm workflow

Executing data operations

Executing pipeline operations

Running maintenance operations

Further reading

Chapter 8: Creating an End-to-End Machine Learning Workflow

Chapter 8: Creating an End-to-End Machine Learning Workflow

Technical requirements

NLP example overview

Creating repositories and pipelines

Creating an NER pipeline

Retraining an NER model

Further reading

Chapter 9: Distributed Hyperparameter Tuning with Pachyderm

Chapter 9: Distributed Hyperparameter Tuning with Pachyderm

Technical requirements

Reviewing hyperparameter tuning techniques and strategies

Creating a hyperparameter tuning pipeline in Pachyderm

Further reading

Section 3:Pachyderm Clients and Tools

Section 3:Pachyderm Clients and Tools

Chapter 10: Pachyderm Language Clients

Chapter 10: Pachyderm Language Clients

Technical requirements

Using the Pachyderm Go client

Using the Pachyderm Python client

Further reading

Chapter 11: Using Pachyderm Notebooks

Chapter 11: Using Pachyderm Notebooks

Technical requirements

Enabling Pachyderm Notebooks in Pachyderm Hub

Running basic Pachyderm operations in Pachyderm Notebooks

Creating and running an example pipeline in Pachyderm Notebooks

Further reading

Other Books You May Enjoy

Other Books You May Enjoy

Packt is searching for authors like you

Share Your Thoughts

Customer Reviews

5 star

0

4 star

0

3 star

0

2 star

0

1 star

0

Understanding inputs

We described inputs in Chapter 2, Pachyderm Basics, in detail by providing examples. Therefore, in this section, we'll just mention that inputs define the type of your pipeline. You can specify the following types of Pachyderm inputs:

PFS is a generic parameter that defines a standard pipeline and inputs in all multi-input pipelines.
Cross is an input that creates a cross-product of the datums from two input repositories. The resulting output will include all possible combinations of all datums from the input repositories.
Union is an input that adds datums from one repository to the datums in another repository.
Join is an input that matches datums with a specific naming pattern.
Spout is an input that consumes data from a third-party source and adds it to the Pachyderm filesystem for further processing.
Group is an input that combines datums from multiple repositories based on a configured naming pattern.
Cron is a pipeline...