Chapter 1. The Solution: Data Curation at Scale
Integrating data sources isn’t a new challenge. But the challenge has intensified in both importance and difficulty, as the volume and variety of usable data—and enterprises’ ambitious plans for analyzing and applying it—have increased. As a result, trying to meet today’s data integration demands with yesterday’s data integration approaches is impractical.
In this chapter, we look at the three generations of data integration products and how they have evolved, focusing on the new third-generation products that deliver a vital missing layer in the data integration “stack”: data curation at scale. Finally, we look at five key tenets of an effective data curation at scale system.
Three Generations of Data Integration Systems
Data integration systems emerged to enable business analysts to access converged datasets directly for analyses and applications.
First-generation data integration systems—data warehouses—arrived on the scene in the 1990s. Major retailers took the lead, assembling, customer-facing data (e.g., item sales, products, customers) in data stores and mining it to make better purchasing decisions. For example, pet rocks might be out of favor while Barbie dolls might be “in.” With this intelligence, retailers could discount the pet rocks and tie up the Barbie doll factory with a big order. Data warehouses typically paid for themselves within a year through better buying decisions.
First-generation ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access