April 2019
Intermediate to advanced
212 pages
5h 34m
English
By introducing an exploratory epsilon factor, we immediately improve on the simple greedy strategy by guaranteeing that at least some exploration will take place and we won't get stuck pulling the same currently high-valued arms over and over.
One way to compare strategies is by looking at their total reward, meaning the reward we collect over the course of all of the trial runs through the testing cycle.
We see the results of running a random sampler action-selection method in the following screenshot. Each arm gets chosen roughly the same number of times, no matter the results it gets:

The preceding screenshot shows ...
Read now
Unlock full access