Serverless ETL and Analytics with AWS Glue

By : Vishal Pathak, Subramanya Vajiraya, Noritaka Sekiyama, Tomohiro Tanaka, Albert Quiroga, Ishan Gaur

Serverless ETL and Analytics with AWS Glue

By: Vishal Pathak, Subramanya Vajiraya, Noritaka Sekiyama, Tomohiro Tanaka, Albert Quiroga, Ishan Gaur

Overview of this book

Organizations these days have gravitated toward services such as AWS Glue that undertake undifferentiated heavy lifting and provide serverless Spark, enabling you to create and manage data lakes in a serverless fashion. This guide shows you how AWS Glue can be used to solve real-world problems along with helping you learn about data processing, data integration, and building data lakes. Beginning with AWS Glue basics, this book teaches you how to perform various aspects of data analysis such as ad hoc queries, data visualization, and real-time analysis using this service. It also provides a walk-through of CI/CD for AWS Glue and how to shift left on quality using automated regression tests. You’ll find out how data security aspects such as access control, encryption, auditing, and networking are implemented, as well as getting to grips with useful techniques such as picking the right file format, compression, partitioning, and bucketing. As you advance, you’ll discover AWS Glue features such as crawlers, Lake Formation, governed tables, lineage, DataBrew, Glue Studio, and custom connectors. The concluding chapters help you to understand various performance tuning, troubleshooting, and monitoring options. By the end of this AWS book, you’ll be able to create, manage, troubleshoot, and deploy ETL pipelines using AWS Glue.

Preface

Who this book is for

What this book covers

To get the most out of this book

Download the example code files

Download the color images

Conventions used

Get in touch

Share Your Thoughts

Section 1 – Introduction, Concepts, and the Basics of AWS Glue

Free Chapter

Chapter 1: Data Management – Introduction and Concepts

Types of data processing – OLTP and OLAP

Data warehouses and data marts

Data lakes

Data lakehouse

Data mesh

Distributed computing for big data

AWS Glue

Summary

Chapter 2: Introduction to Important AWS Glue Features

Data integration

Integrating data with AWS Glue

Features of AWS Glue

Summary

Chapter 3: Data Ingestion

Technical requirements

Data ingestion from file/object stores

Data ingestion from JDBC data stores

Data ingestion from streaming data sources

Data ingestion from SaaS data stores

Summary

Section 2 – Data Preparation, Management, and Security

Chapter 4: Data Preparation

Technical requirements

Introduction to data preparation

Data preparation using AWS Glue

Selecting the right service/tool

Summary

Chapter 5: Data Layouts

Technical requirements

Why do we need to pay attention to data layout?

Key techniques to optimally storing data

Optimizing the number of files and each file size

Optimizing your storage with Amazon S3

Summary

Further reading

Section 3 – Tuning, Monitoring, Data Lake Common Scenarios, and Interesting Edge Cases

Chapter 11: Monitoring

Defining an SLA for a data platform

Monitoring the SLA of a data platform

Monitoring the components of a data platform

Analyzing usage

Summary

Chapter 12: Tuning, Debugging, and Troubleshooting

Tuning AWS Glue workloads

Troubleshooting and debugging common issues in AWS Glue ETL

Summary

Chapter 13: Data Analysis

Creating Marketplace connections

Creating the CloudFormation stack

The benefit of ad hoc analysis and how a data lake enables it

Creating and updating Hudi tables using Glue

Creating and updating Delta Lake tables using Glue

Inserting data into Lake Formation governed tables

Consuming streaming data using Glue

Glue’s integration with OpenSearch

Cleaning up

Summary

Chapter 14: Machine Learning Integration

Technical requirements

Glue ML transformations

SageMaker integration

Developing ML pipelines with Glue

Summary

Chapter 15: Architecting Data Lakes for Real-World Scenarios and Edge Cases

Technical requirements

Running a highly selective query on a big fact table using AWS Glue

Dealing with Join performance issues with big fact and small dimension tables in ETL workloads

Solving Join problems involving big fact and big dimension tables using AWS Glue

Reducing time on read operations using AWS Glue grouping

Solving S3 eventual consistency problems using AWS Glue

Summary

Why subscribe?

Other Books You May Enjoy

Packt is searching for authors like you

Share Your Thoughts

Customer Reviews

5 star

4 star

3 star

2 star

1 star

Optimizing the number of files and each file size

The number of files and each file size are also related to the performance of your analytic workloads. In particular, the number of files and file sizes are related to the performance of the data retrieval phase by using an analytic engine in your analytic workloads. To understand the relationship between the number of files and the file size and the performance of the data retrieval process by an analytic engine, we’ll look at how the engine generally retrieves data and returns the result as follows.

The basic process of data retrieval and returning a result is firstly getting a list of files, reading each file, processing the contents of the files based on your queries, and then returning the result. In particular, when processing data in Amazon S3, the analytic engine lists objects in your specified S3 bucket, gets objects, reads the contents, then processes and returns the result. When you use an AWS Glue ETL Spark job...

Serverless ETL and Analytics with AWS Glue

By : Vishal Pathak, Subramanya Vajiraya, Noritaka Sekiyama, Tomohiro Tanaka, Albert Quiroga, Ishan Gaur

Serverless ETL and Analytics with AWS Glue

By: Vishal Pathak, Subramanya Vajiraya, Noritaka Sekiyama, Tomohiro Tanaka, Albert Quiroga, Ishan Gaur

Overview of this book

Related Content you might be interested in

Current Title:

Serverless ETL and Analytics with AWS Glue

Optimizing the number of files and each file size