Part II. Production Readiness
Production readiness means a model can handle sustained traffic without surprises. This part examines the operational work that follows the first successful deployment. It opens by explaining how schedulers, device plug-ins, and resource limits shape throughput and utilization of GPUs. Next, the pieces are tied together with scaling policies, rollout strategies, and failure handling. The concluding chapter shows how logs, metrics, and traces reveal latency, accuracy, and cost information. The aim is to keep performance steady and costs under control as demand grows.
In detail, the chapters in this part cover the following aspects:
-
Chapter 3, “Kubernetes and GPUs”, describes how Kubernetes and GPUs can work well together
-
Chapter 4, “Running in Production”, focuses on the optimization of the model/runtime for production workload.
-
Chapter 5, “Model Observability”, explains the specific observability aspects that make model observability slightly different compared to traditional workload observability on Kubernetes.
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access