Getting Started with the SparkSession
(or HiveContext or SQLContext)Spark SQL DependenciesBasics of SchemasColumn MetadataDataFrame APITransformationsMulti-DataFrame TransformationsPlain Old SQL Queries and Interacting with
Metastore Tables (Hive, Iceberg, etc.)Data Representation in DataFrames and DatasetsTungstenDatasetsInteroperability with RDDs, DataFrames, and Local CollectionsCompile-Time Strong TypingEasier Functional (RDD “like”) TransformationsRelational TransformationsMulti-Dataset Relational TransformationsGrouped Operations on DatasetsExtending with User-Defined Functions (UDFs), Aggregate Functions (UDAFs), and ExpressionsCustom ExpressionsCustom EncodersUser-Defined Table FunctionsData Loading and Saving FunctionsDataFrameWriter and DataFrameReaderFormatsSave ModesMergePartitions (Discovery and Writing)Query OptimizerLogical and Physical PlansStatic Optimizer Rules You Should Understand
(and Working Around Them)Understanding Spark Adaptive Query ExecutionCode GenerationLarge Query Plans and Iterative AlgorithmsSparkSession Extensions and Custom OptimizersDebugging Spark SQL QueriesJDBC/ODBC ServerConclusion