Apache Hudi: The Definitive Guide
by Shiyan Xu, Prashant Wason, Bhavani Sudha Saktheeswaran, Rebecca Bilbro
Chapter 9. Running Hudi in Production
Moving from development to production often brings a new set of operational challenges. This chapter will equip you with the tools and best practices to manage Apache Hudi deployments smoothly in complex environments, ensuring reliable pipelines with minimal overhead.
First, we will explore tools for table management and recovery. You will learn to master the Hudi CLI, a versatile tool for performing routine maintenance, inspecting table metadata for troubleshooting, and executing various operational tasks without writing custom code. We will also cover Hudi’s savepoint and restore operations, which are essential for disaster recovery.
Next, we will focus on integrating Hudi into data platforms. We will cover platform features such as post-commit callbacks, which can be used to trigger downstream processes in messaging systems like Apache Kafka or Apache Pulsar, when actions complete on Hudi tables. You will learn how to set up monitoring and export key metrics to systems like Prometheus or Amazon CloudWatch to maintain visibility into system health. We will also tackle the challenge of metadata consistency, explaining how to use Hudi’s catalog sync services to keep your tables registered and accessible across multiple data catalogs like AWS Glue and DataHub, as well as data warehouses like Google BigQuery, Snowflake, and AWS Redshift.
Finally, we will delve into performance tuning, providing practical advice and proven strategies to optimize ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access