December 2020
Intermediate to advanced
472 pages
13h 36m
English
A
A2C (advantage actor-critic), here
optimize-model logic, here
PPO and, here
train logic, here
A3C (asynchronous advantage actor-critic) algorithm
actor workers, here
history of,
non-blocking model updates, here
n-step estimates, here
overview, here
absorbing state, here
accumulating trace, here
action-advantage function, here, here,
dueling network architecture and, here
optimal, here
actions, here
environments and, here
in examples of RL problems, here
MDPs and, here
overview, here
selection by agents, here
states and, here
variables and, here
action space,
action-value function, here,
actor-critic agents, here
actor-critic methods, here
defined,
multi-agent reinforcement learning ...
Read now
Unlock full access