Summary
Q-learning is an algorithm designed to solve an MDP; that is, a type of control problem that seeks to optimize a variable within a set of constraints. An MDP is built on a Markov chain; a state model in which determining the probability distribution of reaching future states does not require knowledge of any previous states beyond the current one.
An MDP builds on a Markov chain by introducing actions and rewards that can be taken by a learning agent, and allows for choice and decision-making in a stochastic process. Q-learning, as well as other RL algorithms, models the state space of an MDP and progressively reaches an optimal solution by simulating the decisions of a learning agent working within the constraints of the model.
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access