Preface
Today, you can take your phone out in a museum, snap a picture of a painting, and ask a model about the influences the artist drew on and what the piece might be trying to convey. The same model can watch the videos on your phone and give you quick summaries to help you find them later. Vision language models (VLMs) make all of this possible by connecting visual perception and language. They have moved quickly from research prototypes to real products that people use every day.
But building new things with these models is harder than the user experience suggests. The field moves fast, new articles come out daily, and practical guidance is scattered across blog posts, library docs, and informal knowledge passed around at networking events. If you want to use, train, or fine-tune a VLM, it is not obvious how to choose the right architecture, how to curate your datasets, or how to deploy efficiently. You end up piecing the knowledge together yourself.
This book is our attempt to change that. It is the book we wished we had when multimodal work stopped being a research curiosity and became an engineering problem.
We wrote it as a team that has spent years building, documenting, and shipping open source multimodal systems at Hugging Face. Between us we have trained and released VLMs like SmolVLM, integrated dozens of multimodal models into the open source ecosystem, built tooling and demos that make these models accessible to practitioners, and written extensively about the ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access