Chapter 4. Streaming Data Products
In a streaming data mesh, domains own their data. This creates a decentralized data platform to help resolve the issues relating to agility and scalability in the data lake and warehouses. Domains now have to serve other domains their data. So it’s important that they treat their data as products with high quality and trust.
Currently, data engineers are very used to the idea that all their data is in a central data store like a data lake or warehouse. They are used to finding ways to “boil the ocean” (in this case, lake) when working with data. A streaming data mesh allows us to evaporate that idea. In this chapter, we will outline the requirements for streaming data products.
In our careers as data engineers, we have found ourselves writing many wrappers for Apache Spark, a widely used analytics engine for large-scale data processing. Only in the past few years did we fully understand why companies asked us to do this.
Big data tools like Apache Spark, Apache Flink, and Apache Kafka Streams were inaccessible to many engineers who were tasked to solve big data problems. Referring back to Chapter 1, breaking up the monolithic role of a data engineer is a side effect of a data mesh.
This is a very important point because a second side effect is making complex data engineering tools like Spark, Flink, and Kafka Streams more accessible to generalist engineers so they can solve their big data problems. It’s the reason these companies asked us to ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access