Skip to Content
机器学习实战:基于Scikit-Learn、Keras 和TensorFlow (原书第2 版)
book

机器学习实战:基于Scikit-Learn、Keras 和TensorFlow (原书第2 版)

by Aurélien Géron
October 2020
Intermediate to advanced
693 pages
16h 26m
Chinese
China Machine Press
Content preview from 机器学习实战:基于Scikit-Learn、Keras 和TensorFlow (原书第2 版)
强化学习
|
531
我们现在有了一个神经网络策略,该策略将进行观察并输出动作概率。但是我们如何训
练呢?
18.5 评估动作:信用分配问题
如果我们在每一个步骤都知道最佳动作是什么,那我们可以按照通常的方法来训练神经
网络,方法是使估计的概率分布与目标概率分布之间的交叉熵最小。这是常规的有监督
学习。但是在“强化学习”中,智能体得到的唯一指导是通过奖励,而奖励通常是稀疏
的和延迟的。例如,如果智能体使杆子平衡了 100 个步骤,那么如何知道执行的 100 个
动作中哪个是好的,哪些是不好的呢? 它所知道的是,杆子在最后一个动作之后掉下
了,但可以肯定的是,这个最后的动作并不完全负责。这称为信用分配问题:当智能体
得到报酬时,很难知道应该归功于哪些动作(或归咎于哪些动作)。想想看,一只狗在
表现良好之后几小时得到了奖励,它会理解因为什么而获得回报吗?
为了解决这个问题,一种常见的策略是基于动作后获得的所有奖励的总和来评估一个动
作,通常在每个步骤中应用一个折扣因子
γ
(gamma)。折扣后回报的总和称为动作回报。
考虑图 18-6 中的示例。如果一个智能体决定连续三次向右走,并且在第一步之后获得 +10
奖励,在第二步之后获得 0,最后在第三步之后获得 –50,假设我们使用折扣因子
γ
= 0.8,
则第一个动作将得到 10 +
γ
×
0+
γ
2
×
(
-
50)=
-
22 的回报。如果折扣因子接近于 0,则与立
即回报相比,未来回报将不起作用。相反,如果折扣因子接近 1,则远期回报将几乎等
于立即回报。典型的折扣因子从
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

算法技术手册(原书第2 版)

算法技术手册(原书第2 版)

George T.Heineman, Gary Pollice, Stanley Selkow
管理Kubernetes

管理Kubernetes

Brendan Burns, Craig Tracey

Publisher Resources

ISBN: 9787111665977