O'Reilly logo

OpenStack Sahara Essentials by Omar Khedher

Stay ahead with the world's most comprehensive technology and business learning platform.

With Safari, you learn the way you learn best. Get unlimited access to videos, live online training, learning paths, books, tutorials, and more.

Start Free Trial

No credit card required

Boosting Elastic Data Processing performance

Processing data in Hadoop can be very expensive in terms of network latency cost. In a Hadoop cluster, all data is being distributed to all nodes residing in the cluster. Ideally, HDFS splits the data file into chunks to be analyzed by several nodes. Additionally, each chunk of data will be replicated across different machines for data loss resiliency. Basically, every chunk of data is treated in Hadoop as a record broken into a specific format depending on the application logic.

When it comes to processing an assigned record set of data by each node, the Hadoop framework maps each process to the location of data based on the knowledge for the HDFS. To avoid any unneeded network transfers, processes ...

With Safari, you learn the way you learn best. Get unlimited access to videos, live online training, learning paths, books, interactive tutorials, and more.

Start Free Trial

No credit card required