August 2026
Intermediate
312 pages
9h 21m
English
In the RLHF process, the reinforcement learning algorithm slowly updates the model’s weights with respect to feedback from a reward model. The policy—the model being trained—generates completions to prompts in the training set, then the reward model scores them, and the reinforcement learning optimizer takes gradient steps based on this information (see figure 6.1 for an overview). This chapter explains the mathematics and tradeoffs across ...
Read now
Unlock full access