Sign In Start Free Trial
Account

Add to playlist

Create a Playlist

Modal Close icon
You need to login to use this feature.
  • Book Overview & Buying Engineering Lakehouses with Open Table Formats
  • Table Of Contents Toc
Engineering Lakehouses with Open Table Formats

Engineering Lakehouses with Open Table Formats

By : Dipankar Mazumdar, Vinoth Govindarajan
close
close
Engineering Lakehouses with Open Table Formats

Engineering Lakehouses with Open Table Formats

By: Dipankar Mazumdar, Vinoth Govindarajan

Overview of this book

Engineering Lakehouses with Open Table Formats provides detailed insights into lakehouse concepts, and dives deep into the practical implementation of open table formats such as Apache Iceberg, Apache Hudi, and Delta Lake. You’ll explore the internals of a table format and learn in detail about the transactional capabilities of lakehouses. You’ll also get hands on with each table format with exercises using popular computing engines, such as Apache Spark, Flink, Trino, and Python-based tools. The book addresses advanced topics, including performance optimization techniques and interoperability among different formats, equipping you to build production-ready lakehouses. With step-by-step explanations, you’ll get to grips with the key components of lakehouse architecture and learn how to build, maintain, and optimize them. By the end of this book, you’ll be proficient in evaluating and implementing open table formats, optimizing lakehouse performance, and applying these concepts to real-world scenarios, ensuring you make informed decisions in selecting the right architecture for your organization’s data needs.
Table of Contents (15 chapters)
close
close
13
Other Books You May Enjoy
14
Index

Open Data Lakehouse: A New Architectural Paradigm

Processing large volumes of structured and unstructured data is essential for generating insights and making informed decisions. Over the past few decades, the growing need for both real-time transactional processing and large-scale analytical capabilities has influenced the design of data management systems. Initially, organizations relied on two distinct systems: online transaction processing (OLTP) for high-throughput transactional workloads, and online analytical processing (OLAP) for historical trend analysis and complex querying.

However, as data volumes grew beyond the capacity of traditional data warehouses, particularly due to the rise of semi-structured and unstructured data, a new architectural model emerged: the data lake. Built on low-cost cloud or distributed storage, data lakes decoupled compute from storage and allowed organizations to ingest and store raw data of all types at scale. While this architecture addressed scalability and schema flexibility, it lacked the transactional guarantees and governance features necessary for reliable analytics.

These limitations eventually led to the rise of the lakehouse architecture, a unification of the best features of both data warehouses and data lakes. Lakehouses offer the scalability and openness of data lakes with the data reliability and query performance traditionally associated with data warehouses.

In this chapter, we will cover the following topics:

  • The evolution of data systems
  • Emergence of data lakes as centralized storage for diverse data
  • An introduction to the data lakehouse and its architecture
  • Key attributes that define an open data lakehouse

By the end of this chapter, you’ll have a clear understanding of how data management has evolved from OLTP and OLAP systems to the lakehouse architecture. You will also have gained insight into the core components and key attributes that make an open data lakehouse a powerful solution for modern data needs.

Visually different images
CONTINUE READING
83
Tech Concepts
36
Programming languages
73
Tech Tools
Icon Unlimited access to the largest independent learning library in tech of over 8,000 expert-authored tech books and videos.
Icon Innovative learning tools, including AI book assistants, code context explainers, and text-to-speech.
Icon 50+ new titles added per month and exclusive early access to books as they are being written.
Engineering Lakehouses with Open Table Formats
notes
bookmark Notes and Bookmarks search Search in title playlist Add to playlist download Download options font-size Font size

Change the font size

margin-width Margin width

Change margin width

day-mode Day/Sepia/Night Modes

Change background colour

Close icon Search
Country selected

Close icon Your notes and bookmarks

Confirmation

Modal Close icon
claim successful

Buy this book with your credits?

Modal Close icon
Are you sure you want to buy this book with one of your credits?
Close
YES, BUY

Submit Your Feedback

Modal Close icon
Modal Close icon
Modal Close icon