Skip to Content
机器学习实战:基于Scikit-Learn、Keras 和TensorFlow (原书第2 版)
book

机器学习实战:基于Scikit-Learn、Keras 和TensorFlow (原书第2 版)

by Aurélien Géron
October 2020
Intermediate to advanced
693 pages
16h 26m
Chinese
China Machine Press
Content preview from 机器学习实战:基于Scikit-Learn、Keras 和TensorFlow (原书第2 版)
524
|
第
18
章
18.1 学习优化奖励
在强化学习中,软件智能体在环境中进行观察并采取行动,作为回报,它会获得奖励。
它的目标是学会以一种可以随时间推移最大化其预期回报的方式来采取行动。如果你不
介意拟人化,则可以把正面奖励视为愉悦,把负面奖励视为痛苦(在这种情况下,“奖
励”一词有点误导)。简而言之,该智能体在环境中行动,并通过反复试错来学习,以
最大限度地提高其愉悦并最大限度地减少其痛苦。
这是一个相当广泛的设定,可以应用于各种各样的任务。以下是一些示例(见图 18-1):
a. 该智能体可以是控制机器人的程序。在这种情况下,环境就是现实世界,智能体通
过一组传感器(例如摄像头和触摸传感器)来观察环境,其动作包括发送信号以激
活电动马达。它可能被编程为在到达目的地时获得正奖励,而在浪费时间或走错方
向时获得负奖励。
b. 该智能体可以是控制 Ms.Pac-Man 的程序。在这种情况下,环境是 Atari 游戏的模
拟,动作是 9 个可能的操纵杆位置(左上、下、中心等),观察结果是屏幕截图,而
奖励是游戏点数。
c. 类似地,智能体可以是玩棋盘游戏(例如围棋)的程序。
d. 智能体不必控制物理(或虚拟)移动的事物。例如,它可以是一个智能恒温器,只
要温度接近目标温度并节省能源,它就会获得正回报;而当人们需要调节温度时,
它就会获得负回报,因此智能体商必须学会预测人类的需求。
e. 智能体可以观察股市价格并决定每秒要买卖多少。奖励显然是金钱的盈利或者损失。
图 18-1 :强化学习示例:(a)机器人技术;(
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

算法技术手册(原书第2 版)

算法技术手册(原书第2 版)

George T.Heineman, Gary Pollice, Stanley Selkow
管理Kubernetes

管理Kubernetes

Brendan Burns, Craig Tracey

Publisher Resources

ISBN: 9787111665977