Chapter 15. Serving the Expert
You now have a private expert model that you can run! Over the past 14 chapters, you took a base Gemma 4 model, curated a medical Q&A dataset, trained a LoRA adapter, aligned it with DPO, merged it, and compressed it into a single GGUF file that loads on a laptop and still answers medical questions well. That file is the payoff of the whole book. But a file on disk is not a product. Nothing in your existing software can call it, no user can reach it, and it does nothing at all until a process loads it into memory and listens for requests.
That is just as true for the image models you built earlier. Back in Chapters 12 and 13, you left the language world entirely and fine-tuned diffusion models, teaching one a specific subject with DreamBooth and another a specific artistic style with a style LoRA.
All these models have the same problem: a trained checkpoint is not a service until something serves it. That is what this chapter is going to solve.
The serving principles ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access