Chapter 22. Data Validation Is More Than Summary Statistics
Emily Riederer
Which of these numbers doesn’t belong? –1, 0, 1, NA.
It may be hard to tell. If the data in question should be non-negative, –1 is clearly wrong; if it should be complete, the NA is problematic; if it represents the signs to be used in summation, 0 is questionable. In short, there is no data quality without data context.
Data-quality management is widely recognized as a critical component of data engineering. However, while the need for always-on validation is uncontroversial, approaches vary widely. Too often, these approaches rely solely on summary statistics or basic, univariate anomaly-detection methods that are easily automated and widely scalable. However, in the long run, context-free data-quality checks ignore the nuance and help us detect more-pernicious errors that may go undetected by downstream users.
Defining context-enriched business rules as checks on data quality can complement statistical approaches to data validation by encoding domain knowledge. Instead of just defining high-level requirements (e.g., “non-null”), we can define expected interactions between different fields in our data (e.g., “lifetime payments are less than lifetime purchases for each ecommerce customer”).
This enables the exploration of internal consistency across fields in one or more datasets—not just the reasonableness ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access