Skip to Content
Real-Time Analytics: Techniques to Analyze and Visualize Streaming Data
book

Real-Time Analytics: Techniques to Analyze and Visualize Streaming Data

by Byron Ellis
July 2014
Beginner to intermediate
432 pages
10h 54m
English
Wiley
Content preview from Real-Time Analytics: Techniques to Analyze and Visualize Streaming Data

Chapter 5Processing Streaming Data

Now that data is flowing through a data collection system, it must be processed. The original use case for both Kafka and Flume specified Hadoop as the processing system. Hadoop is, of course, a batch system. Although it is very good at what it does, it is hard to achieve processing rates with latencies shorter than about 5 minutes.

The primary source of this limit on the rate of batch processing is startup and shutdown cost. When a Hadoop job starts, a set of input splits is first obtained from the input source (usually the Hadoop Distributed File System, known as HDFS, but potentially other locations). Input splits are parceled into separate mapper tasks by the Job Tracker, which may involve starting new virtual machine instances on the worker nodes. Then there is the shuffle, sort, and reduce phase.

Although each of these steps is fairly small, they add up. A typical job start time requires somewhere between 10 and 30 seconds of “wall time,” depending on the nature of the cluster. Hadoop 2 actually adds more time to the total because it needs to spin up an Application Manager to manage the job. For a batch job that is going to run for 30 minutes or an hour, this startup time is negligible and can be completely ignored for performance tuning. For a job that is running every 5 minutes, 30 seconds of start time represents a 10 percent loss of performance.

Real-time processing frameworks, the subject of this chapter, get around this setup and ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

Large-scale Real-time Stream Processing and Analytics

Large-scale Real-time Stream Processing and Analytics

O'Reilly Media, Inc.
Practical Real-time Data Processing and Analytics

Practical Real-time Data Processing and Analytics

Shilpi Saxena, Selva raj Ramasamy, Prateek Bhati, Saurabh Gupta

Publisher Resources

ISBN: 9781118838020Purchase book