July 2016
Beginner to intermediate
506 pages
11h 23m
English
In Chapter 4, Hadoop and MapReduce Framework for R, you learned about Hadoop and MapReduce frameworks that enable users to process and analyze massive datasets stored in the Hadoop Distributed File System (HDFS). We launched a multi-node Hadoop cluster to run some heavy data crunching jobs using R language which would not be otherwise achievable on an average personal computer with any of the R distributions installed. We also said that although Hadoop is extremely powerful, it is generally recommended for data that greatly exceeds the memory limitations due to its rather slow processing. In this chapter we would like to present Apache Spark engine–a faster way to process and analyze Big Data. After ...
Read now
Unlock full access