book

Vision Language Models

by Merve Noyan, Andrés Marafioti, Miquel Farré, Orr Zohar

July 2026

Intermediate to advanced

300 pages

9h 30m

English

O'Reilly Media, Inc.

Read now

Unlock full access

Brief Table of Contents (Not Yet Final)
1. Introduction to Vision and Language
Brief Introduction to Computer VisionSignal Decomposition TechniquesFilters and Feature extraction kernelsTransformers and its origins in languageVision TransformersBrief Introduction to Hugging Face Open-Source EcosystemConclusion
2. Vision Language Model Applications
Image CaptioningBackgroundImage Captioning Evaluation Metrics.Visual Question Answering (VQA)Feature-Based approachesAttention Mechanism-based approachesTransformer-based modelsVisual ReasoningVisual Language RetrievalHow VLR worksDocument UnderstandingDocument RetrievalInformation extractionLooking aheadVideo UnderstandingInstance Localization with VLMsZero-shot Object DetectionImage-guided Detection
3. Core Architectures of Vision-Language Models
The Key to Combining Information: Multimodal AttentionSelf-Attention: Finding Relationships within a SequenceCross-Attention: Bridging Two Different StreamsModern VLM blueprints: Connecting Pre-trained GiantsThe Adapter approach: Cross-Attention (e.g., Flamingo)The “Unified Sequence” approach: Self-Attention (e.g., SmolVLM)Comparing Architectures: Which Way to Go?Foundational concepts in VLM DesignThe fusion framework: Early or Late?Early Fusion: Integrating from the Ground UpLate Fusion: Independent Strengths, Unified DecisionsConnecting the Spectrum to Modern ArchitecturesThe Classic: Encoder-Decoder FrameworksConclusion: From Blueprints to a New Engineering RealityEnd-of-Chapter Questions (5–7)
4. Training Data and Preprocessing for VLMs
Looking at the DataImage-Text DatasetsVideo-Text DatasetsVision-Language-Action datasetsBuilding a DatasetData Sourcing at ScaleData Filtering at ScaleSample Diversity at ScaleData Annotation and Quality Validation at ScalePreparing the Dataset for ConsumptionDataset MixturesMixture Ingredients and ProportionsTask-Driven Mixture DesignAblations and EvaluationSummaryQuiz
5. Model Training and Optimization
A bird-eye view of training VLMsThe “How”: What different training paradigms do?The “When”: A Model’s training stagesTraining Vision Language ModelsFirst things first: training dataI heard something about an architect?Loading more than one sample at a timeTraining with a real batch sizeInferring with our trained modelMaking inference faster with KV cacheDealing with high resolution imagesSummaryEnd-of-Chapter Questions
6. Post Training Vision Language Models
Supervised Fine-tuningParameter Efficient Fine-tuningTraining with quantizationIntroduction to TRLMultimodal AlignmentReinforcement Learning from Human FeedbackDirect Preference Optimization and Mixed Preference OptimizationGroup Relative Policy OptimizationSummaryQuiz
7. Deploying Models for Inference at Scale
Inference optimization for VLMsUnderstanding VLM inferenceUnderstanding the KV cacheAttention Optimizations: FlashAttention and BeyondQuantization for VLMsKnowledge Distillation for VLMsExporting Models to Different RuntimesPackaging & Deploying in Real EnvironmentsIt runsEfficient deployment with vLLMProduction optimizationsOn‑device/Edge DeploymentThe Edge LandscapeMLX on Apple SiliconLlama.cppMobile DeploymentPEFT Adapters for Edge CustomizationHybrid PatternsSummaryQuiz
8. Video-Language Models
Foundations of Video UnderstandingKey Differences Between Image and Video ModelsRole of Pre-trained Vision ModelsHistorical Evolution: From 2D Vision to Spatiotemporal ModelingEarly Approaches: Frame-Based and Two-Stream Models3D Convolutional Neural NetworksTransformers for VideoCombining video and language understandingTemporal ModelingChallenges in Temporal ModelingAttention in videoVideo-language datasetsPretrainingPost-training video-text datasetsKey Models in Video-Language ModelingEncoder ModelsLarge Multimodal ModelsRunning InferenceTraining and Fine-tuning video LMMsEfficiency in Video Language ModelsToken efficiencyTraining efficiencyEnd-of-Chapter Questions
9. Document AI
Introduction to Document AIInformation ExtractionDocument ParsingMultimodal Document RetrievalApproaches in Solving Document AI ProblemsEarly Document AI Models for information extraction and document classificationDocument AI with Modern Vision Language ModelsFine-tuning Document ModelsFine-tuning KOSMOS2.5 with transformersFine-tuning SmolVLM2 with TRLDocument RAGDocument RetrievalBuilding e2e document RAG pipelineSummary

10. Any to Any Models
Unified Vocabulary ModelsMonolithic architectureFactorized headsHybrid multi-objective modelsVariational autoencoders (VAEs)Connecting continuous latents to language models through diffusionPutting it all together: from prompt to generated outputLate conditioning modelsQuery spaces: how to represent intentConnectors: bridging the gapHow generators use conditioningHands on: MetaQueries and OmniVinci modelsTrainingTask balance and data mixingStaged training: divide and conquerArchitecture specific loss functionsGetting your hands dirtySummaryQuiz
11. Advanced Topics and Cutting-Edge Research
Agentic Vision Language ModelsIntroduction to AgentsIntroduction to SmolagentsComputer Use AgentsVision Language Action modelsClosing the Perception-Action LoopFrom VLM to VLAModel landscape overviewSummaryQuiz
About the Authors

Content preview from Vision Language Models

Chapter 2. Vision Language Model Applications

Vision Language Models (VLMs) are models that can interpret both image and text. VLMs, typically trained on massive vision, vision-language, and language datasets, can accomplish a wide variety of vision tasks like identifying objects, actions, and scenes, as well as understand text for multimodal tasks.

In this chapter, take a look at multimodal tasks and models, how they are employed practically, and how to evaluate them.

Image Captioning

Image captioning is the task of describing the visual content of an image in natural language, see Figure 2-1 to quickly get it. Image captioning sits at the intersection of computer vision and natural language processing. It requires a model to understand an image (identify objects, attributes, and their relationships) and generate a coherent sentence describing the image. This task is a fundamental problem in artificial intelligence that connects vision and language.

Figure 2-1. Image captioning input, ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.

Julian F.

Head of Cybersecurity

I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.

Addison B.

Field Engineer

I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.

Amir M.

Data Platform Tech Lead

I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.

Mark W.

Embedded Software Engineer

Designing Large Language Model Applications

Publisher Resources

ISBN: 9798341624030Errata Page

Cloud Computing

Data Engineering

Data Science

AI & ML

Programming Languages

Software Architecture

IT/Ops

Security

Design