Chapter 12. Diffusion Models
In Chapter 11, you taught a transformer to look at pixels and write words. In this chapter, you’ll meet a different kind of model entirely, one that does the opposite. Diffusion models don’t read images, they paint them. You might be familiar with these types of models already with ChatGPT images, Gemini’s Nano Banana, or Midjourney.
They don’t generate pixels left-to-right, the way a language model generates tokens. They start from pure static and refine it, one careful step at a time, until something coherent that matches your text description emerges. This chapter is the conceptual handoff; you’ll see exactly how the architecture differs from the autoregressive transformers you’ve been working with, why it exists, and, crucially, why the fine-tuning playbook you already know still applies almost unchanged. Same LoRA, same trainer. Different math, different output.
By the end of this chapter, you’ll be able to load Stable Diffusion, generate images from text prompts, ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access