Foreword
When a research field is this young, it is tempting to describe it through its milestones. I’ve had the privilege to have a front-row seat for, and sometimes contribute to, each of this field’s milestones:
The realization that large-scale pretraining enables magic (BigTransfer)
That the exact transformer architecture from language can work for vision (ViT)
I will never forget the night CLIP was published, marvelling at the possibilities training on images and free-form text opens up.
Finally, I’ve witnessed LLaVA, Flamingo, and PaLI be the sparks of the multimodal chatbots that we now have.
At the time, each of these steps felt simultaneously, absolutely unexpected and completely obvious in hindsight.
That is how progress feels when you are close to it.
But milestones are not the same as understanding. A paper can tell you what worked. A model card can tell you what was released. A demo can make the result feel effortless. None of these, by themselves, teaches you how to build a vision language model (VLM), how to debug it when it fails, or how to make the many small decisions that can make or break the final system.
That is where this book comes in. Merve, Miquel, Andi, and Orr do not treat any of this as magic. They open the field up. They show how VLMs are trained, how the data is prepared, what post-training means, and what deployment looks like. But not just that; they also explain how the same ideas and principles extend to documents, videos, any-to-any “omni” ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access