August 2026
Intermediate
312 pages
9h 21m
English
In this chapter, we provide a cursory overview of reinforcement learning from human feedback (RLHF) training before getting into the specifics later in the book. RLHF, while optimizing a simple loss function, involves training multiple different AI models in sequence and then linking them together in a complex online optimization.
Here, we introduce the core objective of RLHF: optimizing a proxy reward for human preferences with a distance-based regularizer (we also show how it relates to classical ...
Read now
Unlock full access