Foundations of Deep Reinforcement Learning: Theory and Practice in Python
by Laura Graesser, Wah Loon Keng
2. REINFORCE
This chapter introduces the first algorithm of the book, REINFORCE.
The REINFORCE algorithm, invented by Ronald J. Williams in 1992 in his paper “Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning” [148], learns a parametrized policy which produces action probabilities from states. Agents use this policy directly to act in an environment.
The key idea is that during learning, actions that resulted in good outcomes should become more probable—these actions are positively reinforced. Conversely, actions which resulted in bad outcomes should become less probable. If learning is successful, over the course of many iterations action probabilities produced by the policy shift to distribution that ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access