Chapter 18. Data Engineering from a Data Scientist’s Perspective
Bill Franks
People have focused on the ingestion and management of data for decades, but only recently has data engineering become a widespread role.1 Why is that? This chapter offers a somewhat contrarian view.
Database Administration, ETL, and Such
Historically, people working with enterprise data focused on three primary areas. First were those who manage raw data collection into source systems. Second were those focused on ETL operations. Until recently, ETL roles were overwhelmingly focused on relational databases. Third were database administrators who manage those relational systems.
The work of these traditional data roles is largely standardized. For example, database administrators don’t tell a database which disks to store data on or how to ensure relational integrity. Since relational technology is mature, many complex tasks are easy. Similarly, ETL tools have adapters for common source systems, functionality to handle common transformation operations, and hooks into common destination repositories. For years, a small number of mature tools interfaced with a small number of mature data repositories. Life was relatively simple!
Why the Need for Data Engineers?
The roles described previously still exist today in their traditional states. However, those roles are no longer sufficient. Data engineers have ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access