Chapter 6. Automate Your Pipeline Tests
Tom White
By sticking to the following guidelines when building data pipelines, and treating data engineering like software engineering, you can write well-factored, reliable, and robust pipelines.
Build an End-to-End Test of the Whole Pipeline
Don’t put any effort into what the pipeline does at this stage. Focus on infrastructure: how to provide known input, do a simple transform, and test that the output is as expected. Use a regular unit-testing framework like JUnit or pytest.
Use a Small Amount of Representative Data
It should be small enough that the test can run in a few minutes at most. Ideally, this data is from your real (production) system (but make sure it is anonymized).
Prefer Textual Data Formats over Binary
Data files should be diff-able, so you can quickly see what’s happening when a test fails. You can check the input and expected outputs into version control and track changes over time.
If the pipeline accepts or produces only binary formats, consider adding support for text in the pipeline itself, or do the necessary conversion in the test.
Ensure That Tests Can Be Run Locally
Running tests locally makes debugging test failures as easy as possible. Use in-process versions of the systems you are using, like Apache Spark’s local mode or Apache HBase’s minicluster, to provide a self-contained local environment.
Minimize ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access