Chapter 50. Modern Metadata for the Modern Data Stack
Prukalpa Sankar
The data world recently converged around the best set of tools for dealing with massive amounts of data, aka the modern data stack. The good? It’s super fast, easy to scale up in seconds, and requires little overhead. The bad? It’s still a noob at bringing governance, trust, and context to data.
That’s where metadata comes in. In the past year, I’ve spoken to more than 350 data leaders to understand their challenges with traditional solutions and construct a vision for modern metadata in the modern data stack. The four characteristics of modern metadata solutions are introduced here.
Data Assets > Tables
Traditional data catalogs were built on the premise that tables were the only asset that needed to be managed. That’s completely different now.
BI dashboards, code snippets, SQL queries, models, features, and Jupyter notebooks are all data assets today. The new generation of metadata management needs to be flexible enough to intelligently store and link different types of data assets in one place.
Complete Data Visibility, Not Piecemeal Solutions
Earlier data catalogs made significant strides in improving data discovery. However, they didn’t give organizations a single source of truth.
Information about data assets is usually spread across different places—data-lineage tools, data-quality tools, data-prep ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access