April 2019
Intermediate to advanced
212 pages
5h 34m
English
A CartPole state-action-reward diagram might look like the following. Whatever state we're in, we can choose from precisely possible actions (left or right) and can go into one of two other states:

Because the true Q-values of each state depend recursively on the Q-values of later states, as these values backpropagate through the network, the Q-value estimates become more and more accurate. As the multidimensional array of Q-values continues to get updated, the agent's actions become more likely to accurately lead it to winning states.
Read now
Unlock full access