Skip to Content
Reinforcement Learning and Stochastic Optimization
book

Reinforcement Learning and Stochastic Optimization

by Warren B. Powell
March 2022
Intermediate to advanced
1136 pages
29h 55m
English
Wiley
Content preview from Reinforcement Learning and Stochastic Optimization

11 Designing Policies

Now that we have learned how to model a sequential decision problem and simulate an exogenous process W1,,Wt,, we return to the challenge of finding a policy that solves our objective function from chapter 9

StartLayout 1st Row max Underscript pi element-of normal upper Pi Endscripts double-struck upper E left-brace sigma-summation Underscript t equals 0 Overscript upper T Endscripts upper C Subscript t Baseline left-parenthesis upper S Subscript t Baseline comma upper X Subscript t Superscript pi Baseline left-parenthesis upper S Subscript t Baseline right-parenthesis right-parenthesis vertical-bar upper S 0 right-brace period EndLayout  (11.1)

objective function has been the basis of our “model first, then solve” approach. But now it is time to solve. This leaves us with the question: How in the world do we search over some arbitrary class of policies?

This is precisely the reason that this form of the objective function is popular with mathematicians who do not care about computation, or in communities where it is already clear what type of policy is being used. However, equation (11.1) is not widely used, and we believe the reason is that there has not been a natural path to computation. In fact, entire fields have emerged which focus on particular classes of policies.

In this chapter, we address the problem of searching over policies in a general way. Our approach is quite practical in that we organize our search using classes of policies that are widely used either in practice or in the research literature. Instead of focusing on a particular hammer looking for a nail, we cover all four classes of policies, with the knowledge that when you settle on an approach, it will come from one of the four classes, or possibly a hybrid of two (or more).

We start by clarifying one area ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

Optimization and Machine Learning

Optimization and Machine Learning

Rachid Chelouah, Patrick Siarry
Machine Learning Design Patterns

Machine Learning Design Patterns

Valliappa Lakshmanan, Sara Robinson, Michael Munn

Publisher Resources

ISBN: 9781119815037Purchase Link