Skip to Content
Reinforcement Learning from Human Feedback, Video Edition
video

Reinforcement Learning from Human Feedback, Video Edition

by Nathan Lambert
August 2026
Intermediate
8h 47m
English
Manning Publications
Closed Captioning available in German, English, Spanish, French, Italian, Japanese

Overview

In Video Editions the narrator reads the book while the content, figures, code listings, diagrams, and text appear on the screen. Like an audiobook that you can also watch as a video.

"A masterful synthesis of the field’s intellectual roots and its practical tools.”
—Saurabh Sawant, Microsoft


Reinforcement Learning from Human Feedback: LLM alignment and post-training helps you understand how modern AI models can be adapted to better match the needs and expectations of their users. Rather than surveying the vast field of reinforcement learning, elite AI researcher Nathan Lambert concentrates exclusively on RLHF and its immediate importance to post-training generative AI models.

This compact book gets right to the point. Early chapters establish the training overview, explain instruction fine-tuning, and build reliable reward models. The middle chapters transition into the heart of alignment, exploring core policy gradient algorithms, Direct Preference Optimization (DPO), and inference-time scaling. Later chapters tackle the messy reality of data, guiding you through preference data collection, synthetic data generation, and the nuances of function calling.

As you go, you will see how these post-training methods actually work, including their unique compute costs and latency trade-offs. You will explore common failure modes, such as qualitative over-optimization, reward hacking, and the unreliability of external evaluation comparisons. Difficult concepts like KL regularization, proximal policy optimization, and generative reward modeling are clarified with hands-on experiments.

Reinforcement Learning from Human Feedback avoids irrelevant academic details in favor of immediate, practical value. Everything author Nathan Lambert includes appears because a modern RLHF project requires it. He skillfully explains complex post-training pipelines by making every detail concrete, connecting isolated abstractions directly to the goal of making models safer, smarter, and perfectly tuned to a desired style.

The book’s seventeen short chapters lay out the core material, while supplements like vocabulary definitions, compute cost management, evaluation variance, and training performance tracking appear in handy appendixes. The result is a logically flowing book that remains highly navigable and technically deep without getting bogged down in unnecessary theory.

About the Technology


About the Book


What's Inside
  • Core RLHF implementations and Direct Alignment Algorithms
  • Building robust preference and synthetic data pipelines
  • Evaluating models and crafting specific AI personas


About the Reader
For established engineers, AI scientists, and students trying to get a practical foothold in AI model alignment.

About the Author
Dr. Nathan Lambert is a leading AI researcher known for leading post-training at the Allen Institute for AI. With previous experience at HuggingFace, DeepMind, and Meta, he is a passionate advocate for open models. His work focuses on increasing access to, and the understanding of, AI technology—empowering readers to contribute to the advancement of AI outside closed corporate labs.

Quotes
The definitive reference and encyclopedia for reinforcement learning.
- Sebastian Raschka, Author of Build a Large Language Model (From Scratch)

The most complete and practical book on RLHF today.
- Andrew Carr, Cartwheel

An essential guide to the modern post-training stack.
- Edward Beeching, Hugging Face

The first complete reference on reinforcement learning, from the fundamentals to the most widely used modern algorithms.
- Sergey Levine, UC Berkeley

Does a fantastic job of distilling years of work into a highly accessible format.
- Yacine Jernite, Hugging Face

Nathan is the right person to educate the next generation of AI practitioners on RLHF and post-training, central components of the modern model-building pipeline.
- Arvind Narayanan, Princeton University

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Watch now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

Creating Online Videos That Engage Viewers

Creating Online Videos That Engage Viewers

Dante M. Pirouz, Allison R. Johnson, Matthew Thomson, Raymond Pirouz

Publisher Resources

ISBN: 9781633434301VEPublisher SupportOtherPublisher WebsitePurchase Link