Foreword by Matei Zaharia
Bringing machine learning and AI into production is not just about the model. Most of the work, and most of the risk, lies in the operational process surrounding it: tracking what you built and how, reproducing a result months later, deploying it safely, and knowing when it has quietly stopped working. As organizations race to bring models, and now agents, into production, that work has only grown more important.
I have spent much of my career on this problem. When we built MLflow, we kept hearing the same thing from teams everywhere: machine learning development is complex in ways that traditional software is not. There are too many tools to stitch together, it is hard to track what went into a result, hard to reproduce it later, and hard to move a model out of a data science notebook and into production. None of that is about the algorithm. We built MLflow to be open, working with any library or tool rather than locking teams into one stack, and later extended that approach to govern data and AI together in Unity Catalog. The goal throughout has been to make ML and AI development as robust and predictable as the best traditional software.
What has been missing is a resource that connects these pieces. There is no shortage of tutorials for any single tool in ML engineering, but practitioners demanded a guide showing how they fit into one coherent practice, from a first notebook to a monitored, governed AI product that a team can trust. That is what this ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access