In this event, we'll examine Spark SQL, a new Alpha component that is part of the Apache Spark 1.0 release. Spark SQL lets developers natively query data stored in both existing RDDs and external sources such as Apache Hive. A key feature of Spark SQL is the ability to blur the lines between relational tables and RDDs, making it easy for developers to intermix SQL commands that query external data with complex analytics. In addition to Spark SQL, we'll explore the Catalyst optimizer framework, which allows Spark SQL to automatically rewrite query plans to execute more efficiently.
- Title: Performing Advanced Analytics on Relational Data with Spark SQL
- Release date: July 2014
- Publisher(s): O'Reilly Media, Inc.
- ISBN: 978149190828
You might also like
Advanced Analytics with Spark, 2nd Edition
In the second edition of this practical book, four Cloudera data scientists present a set of …
Expanded Edition (August 2018) Updated with Design Patterns episodes from the Clean Code series from Clean …
51+ hours of video instruction. Overview The professional programmer’s Deitel® video guide to Python development with …
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow, 2nd Edition
Through a series of recent breakthroughs, deep learning has boosted the entire field of machine learning. …