August 2026
Intermediate
312 pages
9h 21m
English
Reinforcement learning from human feedback, also referred to as reinforcement learning from human preferences in early literature, emerged to optimize machine learning models in domains where specifically designing a reward function is hard. The word preferences is at the center of the RLHF process: human preferences are what we’re trying to model and what fuels the data for training. To understand the scope of the challenge in modeling and measuring human preferences, a broader context is needed in understanding what a preference is, ...
Read now
Unlock full access