Chapter 9. Going Beyond Scala
Working in Spark doesn’t mean limiting yourself to Scala or even limiting yourself to the JVM or languages that Spark explicitly supports. There are many ways to use Spark with different languages, including remote procedure call systems like Spark Connect, JVM interop options like Java Native Interface (JNI), Java Native Access (JNA), or Java Native Runtime (JNR), and prebuilt wrappers over these (like PySpark or SparklyR). This chapter will discuss the performance considerations of using other languages in Spark and how to work with existing libraries effectively. You will learn the performance trade-offs of using different languages and how you can use common Spark accelerators to improve your existing pipelines with minimal changes.
Spark’s language interoperability can be thought of in two groups: one is the worker code inside of your transformations (e.g., the lambdas inside of your maps), and the second is being able to specify the transformations on RDDs/Datasets (e.g., the driver program).
Spark has first-party APIs, or built-in support for APIs, to write driver programs and worker code in R, Python, Scala, and Java. The first-party languages share much of the same design, making it easier to reason about their performance. However, even if support is first-party, it does not mean it will be better. In some cases, third-party bindings have taken interesting work1 to minimize overhead that has not been implemented in the first-party languages, ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access