January 2020
Intermediate to advanced
432 pages
10h 18m
English
In Chapter 8, Policy Gradient Methods, we covered how policy gradient methods can fail and then introduced the TRPO method. Here, we talked about the general strategies TRPO uses to address the failings in PG methods. However, as we have seen, TRPO is quite complex and seeing how it works in code was not much help either. This is the main reason we minimized our discussion of the details when we introduced TRPO and instead waited until we got to this section to tell
the full story in a concise manner.
That said, let's review how policy optimization with TRPO or PPO can do what it does:
Minorize-Maximization MM algorithm: Recall that this is where we find the minimum of an upper bound function by finding ...
Read now
Unlock full access