August 2026
Intermediate
312 pages
9h 21m
English
With the core methods in hand, this part turns to the data that drives RLHF. RLHF is fundamentally a data problem: the quality of a model’s behavior is bounded by the quality of the preferences it learns from. The same holds for post-training gener-ally. These chapters explore why RLHF problems don’t have one perfect solution, what preferences actually are, how data is collected, and how AI-generated synthetic data is increasingly replacing human annotation in modern training pipelines.
Chapter 10 steps back from the technical machinery to motivate the broader context of RLHF and why it is such a crucial problem. It asks questions such as, what are preferences, and why should we expect them to improve models? ...
Read now
Unlock full access