November 2019
Intermediate to advanced
296 pages
7h 52m
English
The simple sum of the reward can be defined as follows:

This is just a total of the rewards obtained from the current state in relation to a specific range of future steps.
In reality, the simple metric is not appropriate as a target to be maximized because it can diverge to infinity when the time step increases. The infinity reward is not working well with the mathematical algorithms. To deal with the situation, it is common to use a discounted total reward instead of reinforcement learning. This type of reward is a formulation aimed to express the uncertainty by using the discounted reward. Specifically, it is described ...
Read now
Unlock full access