Coding our first Spark SQL job
In this section, we will discuss the basics of writing/coding Spark SQL jobs in Scala and Java. Spark SQL exposes the rich DataFrame API (http://spark.apache.org/docs/latest/api/scala/index.html#org.apache.spark.sql.DataFrame) for loading and analyzing datasets in various forms. It not only provides operations for loading/analyzing data from structured formats such as Hive, Parquet, and RDBMS, but also provides flexibility to load data from semistructured formats such as JSON and CSV. In addition to the various explicit operations exposed by the DataFrame API, it also facilitates the execution of SQL queries against the data loaded in the Spark.
Let's move ahead and code our first Spark SQL job in Scala and then we ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access