Dataset joining

In this section, we will cover dataset joining techniques. We will also discuss some of Spark's special features for data joining plus some data joining solutions made easy with Spark.

After this section, we will be able to join data for various machine learning needs.

Dataset joining and its tool – the Spark SQL

In preparing datasets for a machine learning project, we often need to combine data from multiple datasets. For relational tables, the task is to join tables through a primary and foreign key relationship.

Joining two or more datasets together sounds easy, but can be very challenging and time consuming. In SQL, SELECT is the most frequently used command. As an example, the following is a typical SQL code to perform a join: ...

Get Apache Spark Machine Learning Blueprints now with the O’Reilly learning platform.

O’Reilly members experience books, live events, courses curated by job role, and more from O’Reilly and nearly 200 top publishers.