Chapter 1. Overview
Organizations today are bursting at the seams with data, including existing databases, output from applications, and streaming data from ecommerce, social media, apps, and connected devices on the Internet of Things (IoT).
We are all well versed on the data warehouse, which is designed to capture the essence of the business from other enterprise systems—for example, customer relationship management (CRM), inventory, and sales transactions systems—and which allows analysts and business users to gain insight and make important business decisions from that data.
But new technologies, including mobile, social platforms, and IoT, are driving much greater data volumes, higher expectations from users, and a rapid globalization of economies.
Organizations are realizing that traditional technologies can’t meet their new business needs.
As a result, many organizations are turning to scale-out architectures such as data lakes, using Apache Hadoop and other big data technologies. However, despite growing investment in data lakes and big data technology—$150.8 billion in 2017, an increase of 12.4% over 20161—just 14% of organizations report ultimately deploying their big data proof-of-concept (PoC) project into production.2
One reason for this discrepancy is that many organizations do not see a return on their initial investment in big data technology and infrastructure. This is usually because those organizations fail to do data lakes right, falling short when it comes ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access