April 2019
Intermediate to advanced
212 pages
5h 34m
English
For each iteration of the MABP task, we have two arms we can potentially pull. When we first start, we have no knowledge of how often either arm will pay out. Our overall goal is to maximize our long-term payout. We start out in an exploration phase and pull each arm multiple times to see how often it pays out.
For example, take the following Bernoulli (binary) bandit, where the payout is either 0 or 1:
| Trial | Arm | Reward |
| 1 | 1 | 0 |
| 2 | 2 | 1 |
| 3 | 1 | 1 |
| 4 | 2 | 0 |
| 5 | 1 | 1 |
| 6 | 2 | 0 |
| 7 | 1 | 1 |
| 8 | 2 | 1 |
| 9 | 1 | 1 |
| 10 | 2 | 1 |
Effectively, this is pure exploration. We are simply switching between arms here, not yet looking at the results we are getting in deciding which arm to pull next. We are not using any other strategy ...
Read now
Unlock full access