Chapter 56. Preventing the Data Lake Abyss
Scott Haines
Everyone has worked under the wrong assumptions at one point or another in their careers, and in no place have I found this more apparent than when it comes to legacy data and a lot of what ends up in most companies’ data lakes.
The concept of the data lake evolved from the more traditional data warehouse, which was originally envisioned as a means to alleviate the issue of data silos and fragmentation within an organization. The data warehouse achieved this by providing a central store where all data can be accessed, usually through a traditional SQL interface or other business intelligence tools. The data lake takes this concept one step further and allows you to dump all of your data in its raw format (unstructured or structured) into a horizontally scalable massive data store (HDFS/S3) where it can be stored almost indefinitely.
Over the course of many years, what usually starts with the best of intentions can easily turn into a black hole for your company’s most valuable asset as underlying data formats change and render older data unusable. This problem seems to arise from three central issues:
A basic lack of ownership for the team producing a given dataset
A general lack of good etiquette or data hygiene when it comes to preserving backward compatibility with ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access