April 2019
Intermediate to advanced
212 pages
5h 34m
English
One interesting variation on the multi-armed bandit problem is bandits with knapsacks. This formulation of the problem was introduced in 2013 and adds in an aspect of resource consumption by the learning agent. The agent has until its resources run out to maximize its reward output.
The bandits with knapsacks formulation is especially useful in economics in framing concepts such as dynamic pricing, where a seller might offer different prices to a customer based on what each customer would be likely to pay. It has many other applications in large-scale finance problems.
Read now
Unlock full access