June 2018
Intermediate to advanced
318 pages
9h 24m
English
Unlike value iteration, in policy iteration we start with the random policy, then we find the value function of that policy; if the value function is not optimal then we find the new improved policy. We repeat this process until we find the optimal policy.
There are two steps in policy iteration:

The steps involved in the policy iteration are as follows:
Read now
Unlock full access