Chapter 64. Tardy Data
Ariel Shaqed
When collecting time-based data, some data is born late, some achieves lateness, and some has lateness thrust upon it. That makes processing the “latest” data challenging. For instance:
Trace data is usually indexed by start time. Data for an ongoing interval is born late and cannot be generated yet.
Collection systems can work slower in the presence of failures or bursts, achieving lateness of generated data.
Distributed collection systems can delay some data, thrusting lateness upon it.
Lateness occurs at all levels of a collection pipeline. Most collection pipelines are distributed, and late data arrives significantly out of order. Lateness is unavoidable; handling it robustly is essential.
At the same time, providing repeatable queries is desirable for some purposes, and adding late data can directly clash with it. For instance, aggregation must take late data into account.
Common strategies align by the way they store and query late data. Which to choose depends as much on business logic as it does on technical advantages.
Conceptually, the simplest strategy is to update existing data with late data. Each item of data, no matter how late, is inserted according to its timestamp. This can be done in a straightforward manner with many databases. It can be performed with simple data storage. But any scaling is hard; for example, new data ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access