Now that we have identified the source of data and its characteristics and frequency of arrival, next we need to consider the various collection tools available for tapping the live data into the application:
- Apache Flume: Flume is a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of log data. It has a simple and flexible architecture based on streaming data flows. It is robust and fault tolerant with tenable reliability mechanisms and many fail over and recovery mechanisms. It uses a simple extensible data model that allows for online analytic application. (Source: https://flume.apache.org/). The salient features of Flume are:
- It can easily read streaming data, and ...