September 2016
Beginner to intermediate
326 pages
6h 47m
English
Let's take a look at different tools and techniques used in Hadoop and Spark for Big Data analytics.
While the Hadoop platform can be used for both storing and processing the data, Spark can be used for processing only by reading data into memory.
The following is a tabular representation of the tools and techniques used in typical Big Data analytics projects:
|
Tools used |
Techniques used | |
|---|---|---|
|
Data collection |
Apache Flume for real-time data collection and aggregation Apache Sqoop for data import and export from relational data stores and NoSQL databases Apache Kafka for the publish-subscribe messaging system General-purpose tools such as FTP/Copy |
Real-time data capture Export Import Message publishing Data APIs Screen scraping ... |
Read now
Unlock full access