Skip to Content
Serverless ETL and Analytics with AWS Glue
book

Serverless ETL and Analytics with AWS Glue

by Vishal Pathak, Subramanya Vajiraya, Noritaka Sekiyama, Tomohiro Tanaka, Albert Quiroga, Ishan Gaur
August 2022
Intermediate to advanced
434 pages
10h 34m
English
Packt Publishing
Content preview from Serverless ETL and Analytics with AWS Glue

Chapter 6: Data Management

In the previous chapter, you learned how to optimize your data layout to accelerate performance in query engines and manage the data optimally to reduce costs. This is a really important topic, but it is just one aspect of a data lake. As the volume of data increases, a data lake is used by different stakeholders – not only data engineers and software engineers but also data analysts, data scientists, and sales and marketing representatives. Sometimes, the original data is not easy to use for these stakeholders because the raw data may not be structured well. To make business decisions based on data quickly and effectively, it is important to manage, clean up, and enrich the data so that these stakeholders can understand ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Start your free trial

You might also like

PySpark and AWS: Master Big Data with PySpark and AWS

PySpark and AWS: Master Big Data with PySpark and AWS

AI Sciences
AWS Certified Data Engineer Associate Study Guide

AWS Certified Data Engineer Associate Study Guide

Sakti Mishra, Dylan Qu, Anusha Challa

Publisher Resources

ISBN: 9781800564985Supplemental Content