Book Image

Accelerate Model Training with PyTorch 2.X

By : Maicon Melo Alves

Book Image

Accelerate Model Training with PyTorch 2.X

By: Maicon Melo Alves

Overview of this book

Penned by an expert in High-Performance Computing (HPC) with over 25 years of experience, this book is your guide to enhancing the performance of model training using PyTorch, one of the most widely adopted machine learning frameworks. You’ll start by understanding how model complexity impacts training time before discovering distinct levels of performance tuning to expedite the training process. You’ll also learn how to use a new PyTorch feature to compile the model and train it faster, alongside learning how to benefit from specialized libraries to optimize the training process on the CPU. As you progress, you’ll gain insights into building an efficient data pipeline to keep accelerators occupied during the entire training execution and explore strategies for reducing model complexity and adopting mixed precision to minimize computing time and memory consumption. The book will get you acquainted with distributed training and show you how to use PyTorch to harness the computing power of multicore systems and multi-GPU environments available on single or multiple machines. By the end of this book, you’ll be equipped with a suite of techniques, approaches, and strategies to speed up training , so you can focus on what really matters—building stunning models!

Preface

Who this book is for

What this book covers

To get the most out of this book

Download the example code files

Conventions used

Share Your Thoughts

Download a free PDF copy of this book

Free Chapter

Part 1: Paving the Way

Part 1: Paving the Way

Chapter 1: Deconstructing the Training Process

Chapter 1: Deconstructing the Training Process

Technical requirements

Remembering the training process

Understanding the computational burden of the model training phase

Chapter 2: Training Models Faster

Chapter 2: Training Models Faster

Technical requirements

What options do we have?

Modifying the application layer

Modifying the environment layer

Part 2: Going Faster

Part 2: Going Faster

Chapter 3: Compiling the Model

Chapter 3: Compiling the Model

Technical requirements

What do you mean by compiling?

Using the Compile API

How does the Compile API work under the hood?

Chapter 4: Using Specialized Libraries

Chapter 4: Using Specialized Libraries

Technical requirements

Multithreading with OpenMP

Optimizing Intel CPU with IPEX

Chapter 5: Building an Efficient Data Pipeline

Chapter 5: Building an Efficient Data Pipeline

Technical requirements

Why do we need an efficient data pipeline?

Accelerating data loading

Chapter 6: Simplifying the Model

Chapter 6: Simplifying the Model

Technical requirements

Knowing the model simplifying process

Using Microsoft NNI to simplify a model

Chapter 7: Adopting Mixed Precision

Chapter 7: Adopting Mixed Precision

Technical requirements

Remembering numeric precision

Understanding the mixed precision strategy

Part 3: Going Distributed

Part 3: Going Distributed

Chapter 8: Distributed Training at a Glance

Chapter 8: Distributed Training at a Glance

Technical requirements

A first look at distributed training

Learning the fundamentals of parallelism strategies

Distributed training on PyTorch

Chapter 9: Training with Multiple CPUs

Chapter 9: Training with Multiple CPUs

Technical requirements

Why distribute the training on multiple CPUs?

Implementing distributed training on multiple CPUs

Getting faster with Intel oneCCL

Chapter 10: Training with Multiple GPUs

Chapter 10: Training with Multiple GPUs

Technical requirements

Demystifying the multi-GPU environment

Implementing distributed training on multiple GPUs

Chapter 11: Training with Multiple Machines

Chapter 11: Training with Multiple Machines

Technical requirements

What is a computing cluster?

Implementing distributed training on multiple machines

Index

Other Books You May Enjoy

Other Books You May Enjoy

Packt is searching for authors like you

Share Your Thoughts

Download a free PDF copy of this book

Customer Reviews

5 star

0

4 star

0

3 star

0

2 star

0

1 star

0

How does the Compile API work under the hood?

The Compile API is exactly what its name suggests: it is an entry point to access a set of functionalities PyTorch provides to move from eager to graph execution mode. Besides intermediary components and processes, we also have the compiler, which is an entity that’s responsible for getting the final work done. There are half a dozen compilers available, each one specialized in generating optimized code for a given architecture or device.

The following sections describe the steps that are involved in the compiling process and the components that make all this possible.

Compiling workflow and components

At this point, we can imagine that the compiling process is much more complex than calling a single line in our code. To transform an eager model into a compiled model, the Compile API executes three steps, namely graph acquisition, graph lowering, and graph compilation, as depicted in Figure 3.8:

Figure 3.8 – Compiling workflow

...