Chapter 1. Introduction to Data Virtualization and Data Lakes
As long as humankind has existed, knowledge has been spread out across the world. In ancient times, one had to go to Delphi to see the oracle, to Egypt for the construction knowledge necessary to build the pyramids, and to Babylon for the irrigation science to build the Hanging Gardens. If someone had a simple question, such as, “Were the engineering practices used in building the pyramids and Hanging Gardens similar?” no answer could be obtained in less than a decade. The person would have to spend years learning the languages used in Babylon and Egypt in order to converse with the locals who knew the history of these structures; learn enough about the domains to express their questions in such a way that the answers would reveal whether the engineering practices were similar; and then travel to these locations and figure out who were the correct locals to ask.
In modern times, data is no less spread out than it used to be. Important knowledge—capable of answering many of the world’s most important questions—remains dispersed across the entire world. Even within a single organization, data is typically spread across many different locations. Data generally starts off being located where it was generated, and a large amount of energy is needed to overcome the inertia to move it to a new location. The more locations data is generated in, the more locations data is found in.
Although it no longer takes a decade to answer ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access