Chapter 67. The Data Pipeline Is Not About Speed
Rustem Feyzkhanov
Data-processing pipelines used to be about speed. Now we live in a world of public cloud technologies, where any company can provide additional resources in a matter of seconds. This fact changes our perspective on the way data processing pipelines should be constructed.
In practice, we pay equally for using 10 servers for 1 minute and 1 server for 10 minutes. Because of that, the focus has shifted from optimizing execution time to optimizing scalability and parallelization. Let’s imagine a perfect data-processing pipeline: 1,000 jobs get in, are processed in parallel on 1,000 nodes, and then the results are gathered. This would mean that on any scale, the speed of processing doesn’t depend on the number of jobs and is always equal to that of one job.
And today this is possible with public cloud technologies like serverless computing that are becoming more and more popular. They provide a way to launch thousands of processing nodes in parallel. Serverless implementations like AWS Lambda, Microsoft Azure Functions, and Google Cloud Functions enable us to construct scalable data pipelines with little effort. You just need to define libraries and code, and you are good to go. Scalable orchestrators like AWS Step Functions and Azure Logic Apps allow executing hundreds or thousands of jobs in parallel. Services like ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access