Summary
In this chapter, we discussed various transformation and actions provided by the Spark RDD API. We also discussed various real-life problem statements and solved them by using Spark transformation and actions. At the end, we also discussed the persistence/caching provided by Spark for optimizing the performance.
In the next chapter, we will discuss interactive analytics using Spark SQL.
Get Real-Time Big Data Analytics now with the O’Reilly learning platform.
O’Reilly members experience books, live events, courses curated by job role, and more from O’Reilly and nearly 200 top publishers.