Getting DataOps Right
by Andy Palmer, Michael Stonebraker, Nik Bates-Haus, Liam Cleary, Mark Marinelli
Chapter 4. Key Principles of a DataOps Ecosystem
Having worked with dozens of Global 2000 customers on their data/analytics initiatives, I have seen a consistent pattern of key principles of a DataOps ecosystem that is in stark contrast to traditional “single vendor,” “single platform” approaches that are advocated by vendors such as Palantir, Teradata, IBM, Oracle, and others. An open, best of breed approach is more difficult, but also much more effective in the medium and long term; it represents a winning strategy for a chief data officer, chief information officer, and CEO who believe in maximizing the reuse of quality data in the enterprise and avoids the oversimplified trap of writing a massive check to a single vendor with the belief that there will be “one throat to choke.”
There are certain key principles of a DataOps ecosystem that we see at work every day in a large enterprise. A modern DataOps infrastructure/ecosystem should be and do the following:
-
Highly automated
-
Open
-
Take advantage of best of breed tools
-
Use Table(s) In/Table(s) Out protocols
-
Have layered interfaces
-
Track data lineage
-
Feature deterministic, probabilistic, and humanistic data integration
-
Combine both aggregated and federated methods of storage and access
-
Process data in both batch and streaming modes
I’ve provided more detail/thoughts on each of these in this chapter.
Highly Automated
The scale and scope of data in the enterprise has surpassed the ability of bespoke ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access