Chapter 14. Optimization and Quantization
By now you have a fine-tuned model that works. Over the course of this book, you took a base Gemma model, fed it a curated medical Q&A dataset, trained a LoRA adapter, aligned it with DPO, and confirmed that it answers domain questions better than the stock model. That model lives as a base checkpoint plus a small LoRA adapter, and it runs comfortably on the GPU you trained it on.
The trouble is that the GPU you trained on is not the machine you ship on. Similarly with Stable Diffusion and subject/style tuning, you’ve created an adapter to shape its behavior.
This chapter is about closing the gap between “my model trains successfully” and “it deploys effectively.” You’re going to do three things to your model, each independent and each compounding: you’ll merge the adapter back into the base so there’s a single artifact to ship; you’ll quantize the weights from 16 bits down to 4, cutting the model’s footprint significantly with a loss in quality you ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access