Sign In Start Free Trial

Book Overview & Buying
Table Of Contents

Reproducible Data Science with Pachyderm

By : Svetlana Karslioglu

5 (3)

Reproducible Data Science with Pachyderm

5 (3)

By: Svetlana Karslioglu

Overview of this book

Pachyderm is an open source project that enables data scientists to run reproducible data pipelines and scale them to an enterprise level. This book will teach you how to implement Pachyderm to create collaborative data science workflows and reproduce your ML experiments at scale. You’ll begin your journey by exploring the importance of data reproducibility and comparing different data science platforms. Next, you’ll explore how Pachyderm fits into the picture and its significance, followed by learning how to install Pachyderm locally on your computer or a cloud platform of your choice. You’ll then discover the architectural components and Pachyderm's main pipeline principles and concepts. The book demonstrates how to use Pachyderm components to create your first data pipeline and advances to cover common operations involving data, such as uploading data to and from Pachyderm to create more complex pipelines. Based on what you've learned, you'll develop an end-to-end ML workflow, before trying out the hyperparameter tuning technique and the different supported Pachyderm language clients. Finally, you’ll learn how to use a SaaS version of Pachyderm with Pachyderm Notebooks. By the end of this book, you will learn all aspects of running your data pipelines in Pachyderm and manage them on a day-to-day basis.

Preface

Preface

Who this book is for

What this book covers

To get the most out of this book

Download the example code files

Download the color images

Conventions used

Get in touch

Share Your Thoughts

Section 1: Introduction to Pachyderm and Reproducible Data Science

Section 1: Introduction to Pachyderm and Reproducible Data Science

Free Chapter

Chapter 1: The Problem of Data Reproducibility

Chapter 1: The Problem of Data Reproducibility

Why is reproducibility important?

The reproducibility crisis in science

Demystifying MLOps

Types of data science platforms

Explaining ethical AI

Summary

Further reading

Chapter 2: Pachyderm Basics

Chapter 2: Pachyderm Basics

Reviewing Pachyderm architecture

Learning about version control primitives

Discovering pipeline elements

Summary

Further reading

Chapter 3: Pachyderm Pipeline Specification

Chapter 3: Pachyderm Pipeline Specification

Pipeline specification overview

Understanding inputs

Exploring informational parameters

Exploring transformation

Optimizing your pipeline

Exploring service parameters

Exploring output parameters

Summary

Further reading

Section 2:Getting Started with Pachyderm

Section 2:Getting Started with Pachyderm

Chapter 4: Installing Pachyderm Locally

Chapter 4: Installing Pachyderm Locally

Technical requirements

Installing the required tools

Installing minikube

Installing Docker Desktop

Installing the Pachyderm command-line interface

Enabling autocompletion for Pachyderm

Preparing the Kubernetes environment

Deploying Pachyderm

Accessing the Pachyderm Console

Deleting an existing Pachyderm deployment

Summary

Further reading

Chapter 5: Installing Pachyderm on a Cloud Platform

Chapter 5: Installing Pachyderm on a Cloud Platform

Technical requirements

Installing the required tools

Deploying Pachyderm on Amazon EKS

Deploying the cluster

Deploying Pachyderm on GKE

Deploying the cluster

Deploying Pachyderm on Microsoft AKS

Deploying the cluster

Accessing the Pachyderm console

Summary

Further reading

Chapter 6: Creating Your First Pipeline

Chapter 6: Creating Your First Pipeline

Technical requirements

Pipeline overview

Creating a repository

Creating a pipeline specification

Viewing the pipeline result

Adding another pipeline step

Summary

Further reading

Chapter 7: Pachyderm Operations

Chapter 7: Pachyderm Operations

Technical requirements

Reviewing the standard Pachyderm workflow

Executing data operations

Executing pipeline operations

Running maintenance operations

Summary

Further reading

Chapter 8: Creating an End-to-End Machine Learning Workflow

Chapter 8: Creating an End-to-End Machine Learning Workflow

Technical requirements

NLP example overview

Creating repositories and pipelines

Creating an NER pipeline

Retraining an NER model

Summary

Further reading

Chapter 9: Distributed Hyperparameter Tuning with Pachyderm

Chapter 9: Distributed Hyperparameter Tuning with Pachyderm

Technical requirements

Reviewing hyperparameter tuning techniques and strategies

Creating a hyperparameter tuning pipeline in Pachyderm

Summary

Further reading

Section 3:Pachyderm Clients and Tools

Section 3:Pachyderm Clients and Tools

Chapter 10: Pachyderm Language Clients

Chapter 10: Pachyderm Language Clients

Technical requirements

Using the Pachyderm Go client

Using the Pachyderm Python client

Summary

Further reading

Chapter 11: Using Pachyderm Notebooks

Chapter 11: Using Pachyderm Notebooks

Technical requirements

Enabling Pachyderm Notebooks in Pachyderm Hub

Running basic Pachyderm operations in Pachyderm Notebooks

Creating and running an example pipeline in Pachyderm Notebooks

Summary

Further reading

Why subscribe?

Other Books You May Enjoy

Other Books You May Enjoy

Packt is searching for authors like you

Share Your Thoughts

Exploring output parameters

Output parameters enable you to configure what happens to your processed data after the result lands in the output repository. You can set it up to be placed in an external S3 repository or configure an egress.

s3_out

The s3_out parameter enables your Pachyderm pipeline to write output to an S3 repository instead of the standard pfs/out. This parameter requires a Boolean value. To access the output repository, you would have to use an S3 protocol address, such as s3://<output-repo>. The output repository will still be eponymous to your pipeline's name.

The following code shows how to define an s3_out parameter in YAML format:

s3_out: true

Here's how to do the same in JSON format:

"s3_out": true

Now, let's learn about egress.

egress

The egress parameter enables you to specify an external location for your output data. Pachyderm supports Amazon S3 (the s3:// protocol), Google Cloud Storage (the gs:...

CONTINUE READING

83

Tech Concepts

36

Programming languages

73

Tech Tools

Unlimited access to the largest independent learning library in tech of over 8,000 expert-authored tech books and videos.

Innovative learning tools, including AI book assistants, code context explainers, and text-to-speech.

50+ new titles added per month and exclusive early access to books as they are being written.

Reproducible Data Science with Pachyderm

Search

Your notes and bookmarks