July 2024
Intermediate to advanced
650 pages
17h 23m
English
The book so far has covered most of the popular RL approaches, including the state-of-the-art PPO with its application in Large Language Models for RLHF fine-tuning. You may have noticed that the focus has always been on only one agent in the environment that learns to act optimally using RL training algorithms. However, there is a whole range of settings with more than one agent. These agents in the environment—either individually or in a collaborative manner—try to achieve some goal. A setup involving ...
Read now
Unlock full access