Apache Hudi: The Definitive Guide
by Shiyan Xu, Prashant Wason, Bhavani Sudha Saktheeswaran, Rebecca Bilbro
Chapter 4. Reading from Hudi
The ability to efficiently read and query data is the ultimate purpose of any data lakehouse, directly impacting the speed and flexibility of analytics and machine learning. A deep understanding of Hudi’s read-side capabilities—and how they integrate with various query engines—is therefore paramount for building a performant and reliable data platform. Building upon the foundational concepts of table layouts from Chapter 2 and the write operations we explored in Chapter 3, this chapter combines technical deep dives and practical examples, serving as your definitive guide to reading data from Apache Hudi.
This chapter is organized into three sections to provide a comprehensive exploration of Hudi’s read capabilities. “Integrating with Query Engines” explains how popular query engines interact with Hudi. We will introduce the Hudi read flow, discuss the integration mechanisms for seamless and efficient querying of your Hudi tables, and examine the role of data catalogs in this process.
To ground our discussion in practical application, “Exploring Query Types” showcases Hudi’s diverse read capabilities with examples. We will explore different query types and discuss the related behaviors and configuration options that enable you to solve a wide range of analytical challenges.
The flexibility and power of Hudi’s read operations are enhanced by several important features designed for advanced data lakehouse patterns. “Highlighting Noteworthy Features” ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access