February 2017
Intermediate to advanced
274 pages
5h 58m
English
The introduction of Apache Spark 2.0 is the recent major release of the Apache Spark project based on the key learnings from the last two years of development of the platform:

Source: Apache Spark 2.0: Faster, Easier, and Smarter http://bit.ly/2ap7qd5
The three overriding themes of the Apache Spark 2.0 release surround performance enhancements (via Tungsten Phase 2), the introduction of structured streaming, and unifying Datasets and DataFrames. We will describe the Datasets as they are part of Spark 2.0 even though they are currently only available in Scala and Java.
Refer to the following presentations by key Spark committers ...
Read now
Unlock full access