Foundations of Deep Reinforcement Learning: Theory and Practice in Python
by Laura Graesser, Wah Loon Keng
3. SARSA
In this chapter we look at SARSA, our first value-based algorithm. It was invented by Rummery and Niranjan in their 1994 paper “On-Line Q-Learning Using Connectionist Systems” [118] and was given its name because “you need to know State-Action-Reward-State-Action before performing an update.”1
1. SARSA was not actually called SARSA by Rummery and Niranjan in their 1994 paper “On-Line Q-Learning Using Connectionist Systems” [118]. The authors preferred “Modified Connectionist Q-Learning.” The alternative was suggested by Richard Sutton and it appears that SARSA stuck.
Value-based algorithms evaluate state-action pairs (s, a) by learning one of the value functions—Vπ(s) or Qπ(s, a)—and use these evaluations to select actions. Learning ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access