Chapter 3. Upgrading Spark
When we started writing the second edition of this book, one of the first tasks we had to face was upgrading our examples from Spark 2.2 to Spark 3.3 (and then later 4; it took us a while to finish the book). In our day jobs, we also often face the task of helping people upgrade to new versions of Spark. Upgrading to new versions of Spark is important to be able to take advantage of its many performance improvements; some of these can be as simple as making your code run on the new engine, whereas in other cases, you may need to use newer APIs. In this chapter you will learn about how to identify areas of Spark that have changed and where you may need to update your codebase.
Upgrading to newer versions of Spark is not as simple as bumping the version and basking in the joy of a new engine. While Spark officially aims to follow semantic versioning (SemVer), where it maintains API compatibility within the same major version, it frequently falls down in important places as we’ll explore in this chapter. Notably, “developer APIs”—which high performance users are most likely to take advantage of—are excluded from Spark’s promises around API compatibility. Don’t let these obstacles keep you from upgrading, though; let’s explore how to find the changes that will impact you, not just in 4 but in any version of Spark.
Finding What You Need to Change
We like to classify breaking changes into three types: (1) those that break at compile time (Scala/Java/Kotlin ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access