Practical Real-time Data Processing and Analytics
by Shilpi Saxena, Selva raj Ramasamy, Prateek Bhati, Saurabh Gupta
Accumulators
This is the second method of sharing values/data across different nodes and/or drivers in a Spark job. As is evident from the name of this variable, accumulators are used for counting or accumulating the values. They are the answer to MapReduces counters, and they are different from broadcast variables because they are mutable—the value of an accumulator can change, while the jobs can change the value of an accumulator, but only the driver program can read its value. They work as a great aide for data aggregation and counting across distributed workers of Spark.
Let's assume there is a purchase log for a Walmart store and we need to write a Spark job to detect the count of each type of bad records out of the log. The following ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access