Book Image

Applied Machine Learning and High-Performance Computing on AWS

By : Mani Khanuja, Farooq Sabir, Shreyas Subramanian, Trenton Potgieter

Book Image

Applied Machine Learning and High-Performance Computing on AWS

By: Mani Khanuja, Farooq Sabir, Shreyas Subramanian, Trenton Potgieter

Overview of this book

Machine learning (ML) and high-performance computing (HPC) on AWS run compute-intensive workloads across industries and emerging applications. Its use cases can be linked to various verticals, such as computational fluid dynamics (CFD), genomics, and autonomous vehicles. This book provides end-to-end guidance, starting with HPC concepts for storage and networking. It then progresses to working examples on how to process large datasets using SageMaker Studio and EMR. Next, you’ll learn how to build, train, and deploy large models using distributed training. Later chapters also guide you through deploying models to edge devices using SageMaker and IoT Greengrass, and performance optimization of ML models, for low latency use cases. By the end of this book, you’ll be able to build, train, and deploy your own large-scale ML application, using HPC on AWS, following industry best practices and addressing the key pain points encountered in the application life cycle.

Preface

Who this book is for

What this book covers

To get the most out of this book

Download the example code files

Download the color images

Conventions used

Share Your Thoughts

Download a free PDF copy of this book

Part 1: Introducing High-Performance Computing

Part 1: Introducing High-Performance Computing

Free Chapter

Chapter 1: High-Performance Computing Fundamentals

Chapter 1: High-Performance Computing Fundamentals

Why do we need HPC?

Limitations of on-premises HPC

Benefits of doing HPC on the cloud

Driving innovation across industries with HPC

Further reading

Chapter 2: Data Management and Transfer

Chapter 2: Data Management and Transfer

Importance of data management

Challenges of moving data into the cloud

How to securely transfer large amounts of data into the cloud

AWS online data transfer services

AWS offline data transfer services

Further reading

Chapter 3: Compute and Networking

Chapter 3: Compute and Networking

Introducing the AWS compute ecosystem

Networking on AWS

Selecting the right compute for HPC workloads

Best practices for HPC workloads

Chapter 4: Data Storage

Chapter 4: Data Storage

Technical requirements

AWS services for storing data

Data security and governance

Tiered storage for cost optimization

Choosing the right storage option for HPC workloads

Further reading

Part 2: Applied Modeling

Part 2: Applied Modeling

Chapter 5: Data Analysis

Chapter 5: Data Analysis

Technical requirements

Exploring data analysis methods

Reviewing the AWS services for data analysis

Analyzing large amounts of structured and unstructured data

Processing data at scale on AWS

Chapter 6: Distributed Training of Machine Learning Models

Chapter 6: Distributed Training of Machine Learning Models

Technical requirements

Building ML systems using AWS

Introducing the fundamentals of distributed training

Executing a distributed training workload on AWS

Chapter 7: Deploying Machine Learning Models at Scale

Chapter 7: Deploying Machine Learning Models at Scale

Managed deployment on AWS

Choosing the right deployment option

Batch inference

Real-time inference

Asynchronous inference

The high availability of model endpoints

Blue/green deployments

Chapter 8: Optimizing and Managing Machine Learning Models for Edge Deployment

Chapter 8: Optimizing and Managing Machine Learning Models for Edge Deployment

Technical requirements

Understanding edge computing

Reviewing the key considerations for optimal edge deployments

Designing an architecture for optimal edge deployments

Chapter 9: Performance Optimization for Real-Time Inference

Chapter 9: Performance Optimization for Real-Time Inference

Technical requirements

Reducing the memory footprint of DL models

Key metrics for optimizing models

Choosing the instance type, load testing, and performance tuning for models

Observing the results

Chapter 10: Data Visualization

Chapter 10: Data Visualization

Data visualization using Amazon SageMaker Data Wrangler

Amazon’s graphics-optimized instances

Further reading

Part 3: Driving Innovation Across Industries

Part 3: Driving Innovation Across Industries

Chapter 11: Computational Fluid Dynamics

Chapter 11: Computational Fluid Dynamics

Technical requirements

Introducing CFD

Reviewing best practices for running CFD on AWS

Discussing how ML can be applied to CFD

Chapter 12: Genomics

Chapter 12: Genomics

Technical requirements

Managing large genomics data on AWS

Designing architecture for genomics

Applying ML to genomics

Chapter 13: Autonomous Vehicles

Chapter 13: Autonomous Vehicles

Technical requirements

Introducing AV systems

AWS services supporting AV systems

Designing an architecture for AV systems

ML applied to AV systems

Chapter 14: Numerical Optimization

Chapter 14: Numerical Optimization

Introduction to optimization

Common numerical optimization algorithms

Example use cases of large-scale numerical optimization problems

Numerical optimization using high-performance compute on AWS

Machine learning and numerical optimization

Further reading

Index

Other Books You May Enjoy

Other Books You May Enjoy

Packt is searching for authors like you

Share Your Thoughts

Download a free PDF copy of this book

Customer Reviews

5 star

0

4 star

0

3 star

0

2 star

0

1 star

0

Limitations of on-premises HPC

HPC applications are often based on complex models trained on a large amount of data, which require high-performing hardware such as Graphical Processing Units (GPUs) and software for distributing the workload among different machines. Some applications may need parallel processing while others may require low-latency and high-throughput networking. Similarly, applications such as gaming and video analysis may need performance acceleration using a fast input or output subsystem and GPUs. Catering to all of the different types of HPC applications on-premises might be daunting in terms of cost and maintenance.

Some of the well-known challenges include, but are not limited to, the following:

High upfront capital investment
Long procurement cycles
Maintaining the infrastructure over its life cycle
Technology refreshes
Forecasting the annual budget and capacity requirement

Due to the above-mentioned constraints, planning for an HPC system can be a grueling process, Return On Investment (ROI) for which might be difficult to justify. This can be a barrier to innovation, with slow growth, reduced efficiency, lost opportunities, and limited scalability and elasticity. Let’s understand the impact of each of these in detail.

Barrier to innovation

The constraints of on-premises infrastructure can limit the system design, which will be more focused on the availability of the hardware instead of the business use case. You might not consider some new ideas if they are not supported by the existing infrastructure, thus obstructing your creativity and hindering innovation within the organization.

Reduced efficiency

Once you finish developing the various components of the system, you might have to wait in long prioritized queues to test your jobs, which might take weeks, even if it takes only a few hours to run. On-premises infrastructure is designed to capitalize on the utilization of expensive hardware, often resulting in very convoluted policies for prioritizing the execution of jobs, thus decreasing your productivity and ability to innovate.

Lost opportunities

In order to take full advantage of the latest technology, organizations have to refresh their hardware. Earlier, the typical refresh cycle of three years was enough to stay current, to meet the demands of HPC workloads. However, due to fast technological advancements and a faster pace of innovation, organizations need to refresh their infrastructure more often, otherwise, it might have a larger downstream business impact in terms of revenue. For example, technologies such as Artificial Intelligence (AI), ML, data visualization, risk analysis of financial markets, and so on, are pushing the limits of on-premises infrastructure. Moreover, due to the advent of the cloud, a lot of these technologies are cloud native, and deliver higher performance on large datasets when running in the cloud, especially with workloads that use transient data.

Limited scalability and elasticity

HPC applications rely heavily on infrastructure elements such as containers, GPUs, and serverless technologies, which are not readily available in an on-premises environment, and often have a long procurement and budget approval process. Moreover, maintaining these environments, making sure they are fully utilized, and even upgrading the OS or software packages, requires skills and dedicated resources. Deploying different types of HPC applications on the same hardware is very limiting in terms of scalability and flexibility and does not provide you with the right tools for the job.

Now that we understand the limitations of doing HPC on-premises, let’s see how we can overcome them by running HPC workloads on the cloud.