June 2018
Intermediate to advanced
318 pages
9h 24m
English
We know that in RL environments, we make a transition from one state s to the next state s' by performing some action a and receive a reward r. We save this transition information as
in a buffer called a replay buffer or experience replay. These transitions are called the agent's experience.
The key idea of experience replay is that we train our deep Q network with transitions sampled from the replay buffer instead of training with the last transitions. Agent's experiences are correlated one at a time, so selecting a random batch of training samples from the replay buffer will reduce the correlation between the agent's experience ...
Read now
Unlock full access