Book Image

Applied Machine Learning and High-Performance Computing on AWS

By : Mani Khanuja, Farooq Sabir, Shreyas Subramanian, Trenton Potgieter

Book Image

Applied Machine Learning and High-Performance Computing on AWS

By: Mani Khanuja, Farooq Sabir, Shreyas Subramanian, Trenton Potgieter

Overview of this book

Machine learning (ML) and high-performance computing (HPC) on AWS run compute-intensive workloads across industries and emerging applications. Its use cases can be linked to various verticals, such as computational fluid dynamics (CFD), genomics, and autonomous vehicles. This book provides end-to-end guidance, starting with HPC concepts for storage and networking. It then progresses to working examples on how to process large datasets using SageMaker Studio and EMR. Next, you’ll learn how to build, train, and deploy large models using distributed training. Later chapters also guide you through deploying models to edge devices using SageMaker and IoT Greengrass, and performance optimization of ML models, for low latency use cases. By the end of this book, you’ll be able to build, train, and deploy your own large-scale ML application, using HPC on AWS, following industry best practices and addressing the key pain points encountered in the application life cycle.

Preface

Who this book is for

What this book covers

To get the most out of this book

Download the example code files

Download the color images

Conventions used

Share Your Thoughts

Download a free PDF copy of this book

Part 1: Introducing High-Performance Computing

Part 1: Introducing High-Performance Computing

Free Chapter

Chapter 1: High-Performance Computing Fundamentals

Chapter 1: High-Performance Computing Fundamentals

Why do we need HPC?

Limitations of on-premises HPC

Benefits of doing HPC on the cloud

Driving innovation across industries with HPC

Further reading

Chapter 2: Data Management and Transfer

Chapter 2: Data Management and Transfer

Importance of data management

Challenges of moving data into the cloud

How to securely transfer large amounts of data into the cloud

AWS online data transfer services

AWS offline data transfer services

Further reading

Chapter 3: Compute and Networking

Chapter 3: Compute and Networking

Introducing the AWS compute ecosystem

Networking on AWS

Selecting the right compute for HPC workloads

Best practices for HPC workloads

Chapter 4: Data Storage

Chapter 4: Data Storage

Technical requirements

AWS services for storing data

Data security and governance

Tiered storage for cost optimization

Choosing the right storage option for HPC workloads

Further reading

Part 2: Applied Modeling

Part 2: Applied Modeling

Chapter 5: Data Analysis

Chapter 5: Data Analysis

Technical requirements

Exploring data analysis methods

Reviewing the AWS services for data analysis

Analyzing large amounts of structured and unstructured data

Processing data at scale on AWS

Chapter 6: Distributed Training of Machine Learning Models

Chapter 6: Distributed Training of Machine Learning Models

Technical requirements

Building ML systems using AWS

Introducing the fundamentals of distributed training

Executing a distributed training workload on AWS

Chapter 7: Deploying Machine Learning Models at Scale

Chapter 7: Deploying Machine Learning Models at Scale

Managed deployment on AWS

Choosing the right deployment option

Batch inference

Real-time inference

Asynchronous inference

The high availability of model endpoints

Blue/green deployments

Chapter 8: Optimizing and Managing Machine Learning Models for Edge Deployment

Chapter 8: Optimizing and Managing Machine Learning Models for Edge Deployment

Technical requirements

Understanding edge computing

Reviewing the key considerations for optimal edge deployments

Designing an architecture for optimal edge deployments

Chapter 9: Performance Optimization for Real-Time Inference

Chapter 9: Performance Optimization for Real-Time Inference

Technical requirements

Reducing the memory footprint of DL models

Key metrics for optimizing models

Choosing the instance type, load testing, and performance tuning for models

Observing the results

Chapter 10: Data Visualization

Chapter 10: Data Visualization

Data visualization using Amazon SageMaker Data Wrangler

Amazon’s graphics-optimized instances

Further reading

Part 3: Driving Innovation Across Industries

Part 3: Driving Innovation Across Industries

Chapter 11: Computational Fluid Dynamics

Chapter 11: Computational Fluid Dynamics

Technical requirements

Introducing CFD

Reviewing best practices for running CFD on AWS

Discussing how ML can be applied to CFD

Chapter 12: Genomics

Chapter 12: Genomics

Technical requirements

Managing large genomics data on AWS

Designing architecture for genomics

Applying ML to genomics

Chapter 13: Autonomous Vehicles

Chapter 13: Autonomous Vehicles

Technical requirements

Introducing AV systems

AWS services supporting AV systems

Designing an architecture for AV systems

ML applied to AV systems

Chapter 14: Numerical Optimization

Chapter 14: Numerical Optimization

Introduction to optimization

Common numerical optimization algorithms

Example use cases of large-scale numerical optimization problems

Numerical optimization using high-performance compute on AWS

Machine learning and numerical optimization

Further reading

Index

Other Books You May Enjoy

Other Books You May Enjoy

Packt is searching for authors like you

Share Your Thoughts

Download a free PDF copy of this book

Customer Reviews

5 star

0

4 star

0

3 star

0

2 star

0

1 star

0

Driving innovation across industries with HPC

Every industry and type of HPC application poses different kinds of challenges. The HPC solutions provided by cloud vendors such as AWS help all companies, irrespective of their size, which leads to emerging HPC applications such as reinforcement learning, digital twins, supply chain optimization, and AVs.

Let’s take a look at some of the use cases in life sciences and healthcare, AV, and supply chain optimization.

Life sciences and healthcare

In life sciences and the healthcare domain, a large amount of sensitive and meaningful data is captured almost every minute of the day. Using HPC technology, we can harness this data to gain meaningful insights into critical diseases to save lives by reducing the time taken testing lab samples, drug discovery, and much more, as well as meeting the core security and compliance requirements.

The following are some of the emerging applications in the healthcare and life sciences domain.

Genomics

You can use cloud services provided by AWS to store and share genomic data securely, which helps you build and run predictive or real-time applications to accelerate the journey from genomic data to genomic insights. This helps to reduce data processing times significantly and perform casual analysis of critical diseases such as cancer and Alzheimer’s.

Imaging

Using computer vision and data integration services, you can elevate image analysis and facilitate long-term data retention. For example, by using ML to analyze MRI or X-ray scans, radiology companies can improve operational efficiency and quickly generate lab reports for their patients. Some of the technologies provided by AWS for imaging include Amazon EC2 GPU instances, AWS Batch, AWS ParallelCluster, AWS DataSync, and Amazon SageMaker, which we will discuss in detail in subsequent chapters.

Computational chemistry and structure-based drug design

Combining the state-of-the-art deep learning models for protein classification, advancements in protein structure solutions, and algorithms for describing 3D molecular models with HPC computing resources, allows you to grow and reduce the time to market drastically. For example, in a project performed by Novartis on the AWS cloud, where they were able to screen 10 million compounds against a common cancer target in less than a week, based on their internal calculations, if they had performed a similar experiment in-house, then it would have resulted in about a $40 million investment. By running this experiment on the cloud using AWS services and features, they were able to use their 39 years of computational chemistry data and knowledge. Moreover, it only took 9 hours and $4,232 to conduct the experiment, hence increasing their pace of innovation and experimentation. They were able to successfully identify three out of the ten million compounds, that were screened.

Now that we understand some of the applications in the life sciences and healthcare domain, let us discuss how the automobile and transport industry is using HPC for building AVs.

AVs

The advancement of deep learning models such as reinforcement learning, object detection, and image segmentation, as well technological advancements in compute technology, and deploying models on edge devices, have paved the way for AV. In order to design and build an AV, all the components of the system have to work in tandem, including planning, perception, and control systems. It also requires collecting and processing massive amounts of data and using it to create a feedback loop, for vehicles to adjust their state based on the changing condition of the traffic on roads in real time. This entails having high I/O performance, networking, specialized hardware coprocessors such as GPUs or Field Programmable Gate Arrays (FPGAs), as well as analytics and deep learning frameworks. Moreover, before an AV can even start testing on actual roads, it has to undergo millions of miles of simulation to demonstrate safety performance, due to the high dimensionality of the environment, which is complex and time-consuming. By using the AWS cloud’s virtually unlimited compute and storage capacity, support for advanced deep learning frameworks, and purpose-built services, you can drive a faster time to market. For example, In 2017, Lyft, an American transportation company, launched its AV division. To enhance the performance and safety of its system, it uses petabytes of data collected from its AV fleet to execute millions of simulations every year, which involves a lot of compute power. To run these simulations at a lower cost, they decided to take advantage of unused compute capacity on the AWS cloud, by using Amazon EC2 Spot Instances, which also helped them to increase their capacity to run the simulations at this magnitude.

Next, let us understand supply chain optimization and its processes!

Supply chain optimization

Supply chains are worldwide networks of manufacturers, distributors, suppliers, logistics, and e-commerce retailers that work together to get products from the factory to the customer’s door without delays or damage. By enabling information flow through these networks, you can automate decisions without any human intervention. The key attributes to consider are real-time inventory forecasts, end-to-end visibility, and the ability to track and trace the entire production process with unparalleled efficiency. Your teams will no longer have to handle the minuscular details associated with supply chain decisions. With automation and ML you can resolve bottlenecks in product movement. For example, in the event of a pandemic or natural disaster, you can quickly divert goods to alternative shipping routes without affecting their on-time delivery.

Here are some examples of using ML to improve supply chain processes:

Demand Forecasting: You can combine a time series with additional correlated data, such as holidays, weather, and demographic events, and use deep learning models such as DeepAR to get more accurate results. This will help you to meet variable demand and avoid over-provisioning.
Inventory Management: You can automate inventory management using ML models to determine stock levels and reduce costs by preventing excess inventory. Moreover, you can use ML models for anomaly detection in your supply chain processes, which can help you in optimizing inventory management, and deflect potential issues more proactively, for example, by transferring stock to the right location using optimized routing ahead of time.
Boost Efficiency with Automated Product Quality Inspection: By using computer vision models, you can identify product defects faster with improved consistency and accuracy at an early stage so that customers receive high-quality products in a timely fashion. This reduces the number of customer returns and insurance claims that are filed due to issues in product quality, thus saving costs and time.

All the components of supply chain optimization discussed above need to work together as part of the workflow and therefore require low latency and high throughput in order to meet the goal of delivering an optimal quality product to a customer’s doorstep in a timely fashion. Using cloud services to build the workflow provides you with greater elasticity and scalability at an optimized cost. Moreover, with native purpose-built services, you can eliminate the heavy lifting and reduce the time to market.