August 2026
Intermediate
312 pages
9h 21m
English
A core lesson we learn when using reinforcement learning heavily in our domain is that it is a very strong optimizer, which causes it to pull all the possible increase in reward out of the environment. In modern ML systems, especially with language models, we’re using somewhat contrived notions of environment: the models generate completions (the actions), and an external verifier (that is, a reward model or a scoring function) provides feedback. In this domain, it is common for over-optimization to occur: the RL optimizers ...
Read now
Unlock full access