August 2026
Intermediate
312 pages
9h 21m
English
In this book, we’ve covered many tools for modifying the model to learn from human preferences, verifiable rewards, and other valuable signals. All the methods we use are very powerful and can cause the model to change too much relative to the strong, general model from the previous training stage (often called the reference model). When the model learns too much from a given reward, causing out-of-distribution performance to drop, this is called over-optimization ...
Read now
Unlock full access