Chapter 13. Fine-Tuning Diffusion Models
In Chapter 12, you saw how a diffusion model works from the inside: a frozen VAE, a frozen text encoder, a UNet that predicts noise, and a scheduler that turns predictions into pixels. You also saw, at the conceptual level, where the trainable knobs live. This chapter is where you actually turn and tune those knobs.
You’re going to fine-tune diffusion models for two completely different outcomes. These will be to teach the model a specific subject and a specific style. You’ll use the same model architecture and training loop for both, switching between them by changing only the training data.
The first (subject) teaches the model to draw one specific thing: your dog, your product, your face. The second (style) teaches it a visual vocabulary: your brand’s photography, your illustrator’s hand, an anime line-art look.
By the end of the chapter, you’ll have walked through both recipes end-to-end on a single consumer GPU, you’ll know how to compose multiple ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access