April 2021
Intermediate to advanced
394 pages
10h 11m
English
This is the last chapter of the book. Throughout the book, we have dived deep into many foundational aspects of reinforcement learning (RL). We looked at MDP and at planning in MDP using dynamic planning. We looked at model-free value methods. We talked about scaling up solution techniques using function approximation specifically by using deep learning–based approaches such as DQN. We looked at policy-based methods such as REINFORCE, TRPO, PPO, etc. We unified value and policy optimization methods in the actor-critic (AC) approach. Finally, we looked at how to ...
Read now
Unlock full access