Chapter 4. Spark SQL and DataFrames: Introduction to Built-in Data Sources
In the previous chapter, we explained the evolution of and justification for structure in Spark. In particular, we discussed how the Spark SQL engine provides a unified foundation for the high-level DataFrame and Dataset APIs. Now, we’ll continue our discussion of the DataFrame and explore its interoperability with Spark SQL.
This chapter and the next also explore how Spark SQL interfaces with some of the external components shown in Figure 4-1.
In particular, Spark SQL:
Provides the engine upon which the high-level Structured APIs we explored in Chapter 3 are built.
Can read and write data in a variety of structured formats (e.g., JSON, Hive tables, Parquet, Avro, ORC, CSV).
Lets you query data using JDBC/ODBC connectors from external business intelligence (BI) data sources such as Tableau, Power BI, Talend, or from RDBMSs such as MySQL and PostgreSQL.
Provides a programmatic interface to interact with structured data stored as tables or views in a database from a Spark application
Offers an interactive shell to issue SQL queries on your structured data.
Supports ANSI SQL:2003-compliant commands and HiveQL.
Figure 4-1. Spark SQL connectors and data sources
Let’s begin with how you can use Spark SQL in a Spark application.
Using Spark SQL in Spark Applications
The SparkSession, introduced in Spark 2.0, ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access