Chapter 28. Embrace the Data Lake Architecture
Vinoth Chandar
Oftentimes, data engineers build data pipelines to extract data from external sources, transform it, and enable other parts of the organization to query the resulting datasets. While it’s easier in the short term to just build all of this as a single-stage pipeline, a more thoughtful data architecture is needed to scale this model to thousands of datasets spanning multiple tera/petabytes.
Common Pitfalls
Let’s understand some common pitfalls with the single-stage approach. First of all, it limits scalability since the input data to such a pipeline is obtained by scanning upstream databases—relational database management systems (RDBMSs) or NoSQL stores—that would ultimately stress these systems and even result in outages. Further, accessing such data directly allows for little standardization across pipelines (e.g., standard timestamp, key fields) and increases the risk of data breakages due to lack of schemas/data contracts. Finally, not all data or columns are available in a single place, to freely cross-correlate them for insights or design machine learning models.
Data Lakes
In recent years, the data lake architecture has grown in popularity. In this model, source data is first extracted with little to no transformation into a first set of raw datasets. The goal of these raw datasets is to effectively model an ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access