Chapter 4. Data Journey and Data Storage
This chapter discusses data evolution throughout the lifecycle of a production pipeline. We’ll also look at tools that are available to help manage that process.
As we discussed in the preceding chapters, data is a critical part of the ML lifecycle. As ML data and models change throughout the ML lifecycle, it is important to be able to identify, trace, and reproduce data issues and model changes. As this chapter explains, ML Metadata (MLMD), TensorFlow Metadata (TFMD), and TensorFlow Data Validation (TFDV) are important tools to help you do this. MLMD is a library for recording and retrieving metadata associated with ML workflows, which can help you analyze and debug various parts of an ML system that interact. TFMD provides standard representations of key pieces of metadata used when training ML models, including a schema that describes your expectations for the features in the pipeline’s input data. For example, you can specify the expected type, valency, and range of permissible values in TFMD’s schema format. You can then use a TFMD-defined schema in TFDV to validate your data, using the data validation process discussed in Chapter 2.
Finally, we’ll also introduce some forms of data storage that are particularly relevant to ML, especially for today’s increasingly large datasets such as Common Crawl (380 TiB). In production environments, how you handle your data also determines a large component of your cost structure, the amount of ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access