Chapter 11. Vision Language Models
For the past ten chapters, your work has lived in a world of language and text. Prompts in, completions out. Questions in, answers out. Tokens, like turtles, all the way down. But the real world is more than just language. It’s also visual: pixels, X-rays, dashboards, and factory floors. Vision language models (VLMs) begin to bridge that gap, being able to take an image and a question and produce an answer that demonstrates genuine visual understanding. In this chapter, we’ll explore how they work, using Gemma 4 as our guide. We’ll also explore how to fine-tune a VLM to be a better expert at understanding your specific images. By the end of the chapter, you’ll have a grasp of the two-tower architecture that powers modern VLMs and technology like SigLIP and the Vision Transformer (ViT) architecture that encode images into the same spaces as language tokens. You’ll see a complete vision-tuning pipeline that uses similar technology to what you’ve been using, ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access