October 2015
Beginner to intermediate
254 pages
5h 31m
English
In this chapter, we will cover the following recipes:
In this chapter, we'll be looking at how to bundle our Spark application and deploy it on various distributed environments.
As we discussed earlier in Chapter 3, Loading and Preparing Data – DataFrame the foundation of Spark is the RDD. From a programmer's perspective, the composability of RDDs such as a regular Scala collection is a huge advantage. RDD wraps three vital (and two subsidiary) pieces of information that help in reconstruction of data. This enables fault tolerance. ...
Read now
Unlock full access