The spooling directory source

In an effort to avoid all the assumptions inherent in tailing a file, a new source was devised to keep track of which files have been converted into Flume events and which still need to be processed. The spooling directory source is given a directory to watch for new files to appear. It is assumed that files copied to this directory are complete; otherwise, the source might try and send a partial file. It also assumes that filenames never change; otherwise, the source would loose its place on restarts as to which files have been sent and which have not. The filename condition can be met in log4j by using the DailyRollingFileAppender rather than the RollingFileAppender, however, the currently open file would need to ...

Get Apache Flume: Distributed Log Collection for Hadoop now with the O’Reilly learning platform.

O’Reilly members experience books, live events, courses curated by job role, and more from O’Reilly and nearly 200 top publishers.

Start your free trial

Apache Flume: Distributed Log Collection for Hadoop by Steve Hoffman

The spooling directory source

Don’t leave empty-handed

It’s yours, free.

Check it out now on O’Reilly