Chapter 10: Training with Multiple GPUs | Accelerate Model Training with PyTorch 2.X

Book Overview & Buying
Table Of Contents

Accelerate Model Training with PyTorch 2.X

By : Maicon Melo Alves

4.4 (10)

Buy this Book

Accelerate Model Training with PyTorch 2.X

4.4 (10)

By: Maicon Melo Alves

Buy this Book

Overview of this book

This book, written by an HPC expert with over 25 years of experience, guides you through enhancing model training performance using PyTorch. Here you’ll learn how model complexity impacts training time and discover performance tuning levels to expedite the process, as well as utilize PyTorch features, specialized libraries, and efficient data pipelines to optimize training on CPUs and accelerators. You’ll also reduce model complexity, adopt mixed precision, and harness the power of multicore systems and multi-GPU environments for distributed training. By the end, you'll be equipped with techniques and strategies to speed up training and focus on building stunning models.

Preface

Who this book is for

What this book covers

To get the most out of this book

Download the example code files

Conventions used

Get in touch

Share Your Thoughts

Download a free PDF copy of this book

Free Chapter

Part 1: Paving the Way

Chapter 1: Deconstructing the Training Process

Technical requirements

Remembering the training process

Understanding the computational burden of the model training phase

Quiz time!

Summary

Chapter 2: Training Models Faster

Technical requirements

What options do we have?

Modifying the application layer

Modifying the environment layer

Quiz time!

Summary

Part 2: Going Faster

Chapter 3: Compiling the Model

Technical requirements

What do you mean by compiling?

Using the Compile API

How does the Compile API work under the hood?

Quiz time!

Summary

Chapter 4: Using Specialized Libraries

Technical requirements

Multithreading with OpenMP

Optimizing Intel CPU with IPEX

Quiz time!

Summary

Chapter 5: Building an Efficient Data Pipeline

Technical requirements

Why do we need an efficient data pipeline?

Accelerating data loading

Quiz time!

Summary

Chapter 6: Simplifying the Model

Technical requirements

Knowing the model simplifying process

Using Microsoft NNI to simplify a model

Quiz time!

Summary

Chapter 7: Adopting Mixed Precision

Technical requirements

Remembering numeric precision

Understanding the mixed precision strategy

Enabling AMP

Quiz time!

Summary

Part 3: Going Distributed

Chapter 8: Distributed Training at a Glance

Technical requirements

A first look at distributed training

Learning the fundamentals of parallelism strategies

Distributed training on PyTorch

Quiz time!

Summary

Chapter 9: Training with Multiple CPUs

Technical requirements

Why distribute the training on multiple CPUs?

Implementing distributed training on multiple CPUs

Getting faster with Intel oneCCL

Quiz time!

Summary

Chapter 10: Training with Multiple GPUs

Technical requirements

Demystifying the multi-GPU environment

Implementing distributed training on multiple GPUs

Quiz time!

Summary

Chapter 11: Training with Multiple Machines

Technical requirements

What is a computing cluster?

Implementing distributed training on multiple machines

Quiz time!

Summary

Index

Why subscribe?

Other Books You May Enjoy

Packt is searching for authors like you

Share Your Thoughts

Download a free PDF copy of this book

Accelerate Model Training with PyTorch 2.X

By : Maicon Melo Alves

Accelerate Model Training with PyTorch 2.X

By: Maicon Melo Alves

Overview of this book

Implementing distributed training on multiple GPUs

The NCCL communication backend

Confirmation

Buy this book with your credits?

Submit Your Feedback

Create a Free Account To Continue Reading

Sign in to activate your 7-day free access