August 2016
Intermediate to advanced
420 pages
9h 35m
English
In this chapter, we've introduced some primitives to be able to run distributed jobs on a cluster composed by multiple nodes. We've seen the Hadoop framework and all its components, features, and limitations, and then we illustrated the Spark framework.
In the next chapter, we will dig deep in to Spark, showing how it's possible to do data science in a distributed environment.
Read now
Unlock full access