April 2019
Intermediate to advanced
212 pages
5h 34m
English
In this chapter, we will dive deeper into the topic of multi-armed bandits. We touched on the basics of how they work in Chapter 1, Brushing Up on Reinforcement Learning Concepts, and we'll go over some of the conclusions we reached there. We'll extend our knowledge of the exploration-versus-exploitation process that we learned from our study of Q-learning and apply it to other optimization problems using Q-values and exploration-based strategies.
We will do the following in this chapter:
Read now
Unlock full access