September 2017
Beginner to intermediate
360 pages
8h 13m
English
While working in distributed compute programs and modules, where the code executes on different nodes and/or different workers, a lot of time a need arises to share data across the execution units in the distributed execution setup. Thus Spark has the concept of shared variables. The shared variables are used to share information between the parallel executing tasks across various workers or the tasks and the drivers. Spark supports two types of shared variable:
In the following sections, we will look at these two types of Spark variables, both conceptually and pragmatically.
Read now
Unlock full access