Chapter 4. Running in Production
By now, you have likely deployed your first LLM to run on Kubernetes. It responds to requests, maybe even with decent latency. But production isn’t about working once—it’s about working consistently, at scale and under load.
This chapter is all about that transition. This chapter covers what it takes to make LLM inference stable and efficient in real-world scenarios. This includes expected topics like parameter tuning, along with easily overlooked aspects like runtime memory planning, routing sticky requests to cache-warm replicas, model compression decisions, and advanced topologies that require dedicated network configuration.
Treating a model server like any other container is tempting. Just set a few resource limits, expose a service, and call it a day. But GenAI workloads have unique characteristics (massive models, variable request costs, and GPU-intensive operations) that require specialized configuration. You’ll learn how to configure the platform effectively while avoiding the traps that can quietly erode performance and burn through your GPU budget.
This chapter covers five key areas:
- Model and runtime tuning
-
Selecting, evaluating, compressing, and benchmarking models
- Autoscaling
-
Strategies specific to LLM workloads
- Optimizing vLLM startup time
-
Reducing deployment latency
- LLM-aware routing
-
Intelligent request distribution
- Disaggregated serving
-
Advanced distributed architectures
The most fundamental decision is how to pick ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access