January 2019
Beginner to intermediate
154 pages
4h 31m
English
Spark provides the flexibility to leverage the existing Hive metastore. This will allow users to access table definitions as available to Hive in Spark and to run the same HiveQL in Spark. The difference will be that queries running on Spark will be executed as per the Spark execution plan, and underlying data will be processed as per Spark execution and optimizations. These queries wont follow the MapReduce path, which is the default in Hive.
For many queries, users can see a tremendous performance gain with the Spark execution engine compared to the MapReduce engine, due to the optimized plan of execution in Spark.
Read now
Unlock full access