October 2019
Intermediate to advanced
520 pages
13h 5m
English
Dataproc is GCP's big data-managed service for running Hadoop and Spark clusters. Hadoop and Spark are open source frameworks that handle data processing for big data applications in a distributed manner. Essentially, they provide massive storage for data, whilst also providing enormous processing power to handle concurrent processing tasks.
If we refer back to the End-to-end big data solution section of this chapter, Dataproc is also part of the processing stage. It can be compared to Dataflow; however, Dataproc requires us to provision servers, whereas Dataflow is serverless.
Read now
Unlock full access