Skip to Content
机器学习实战:基于Scikit-Learn、Keras 和TensorFlow (原书第2 版)
book

机器学习实战:基于Scikit-Learn、Keras 和TensorFlow (原书第2 版)

by Aurélien Géron
October 2020
Intermediate to advanced
693 pages
16h 26m
Chinese
China Machine Press
Content preview from 机器学习实战:基于Scikit-Learn、Keras 和TensorFlow (原书第2 版)
训练模型
|
127
改善过拟合模型的一种方法是向其提供更多的训练数据,直到验证误差达到
训练误差为止。
偏差 / 方差权衡
统计学和机器学习的重要理论成果是以下事实:模型的泛化误差可以表示为三个非
常不同的误差之和:
偏差
这部分泛化误差的原因在于错误的假设,比如假设数据是线性的,而实际上是
二次的。高偏差模型最有可能欠拟合训练数据
注 8
。
方差
这部分是由于模型对训练数据的细微变化过于敏感。具有许多自由度的模型
(例如高阶多项式模型)可能具有较高的方差,因此可能过拟合训练数据。
不可避免的误差
这部分误差是因为数据本身的噪声所致。减少这部分误差的唯一方法就是清理
数据(例如修复数据源(如损坏的传感器),或者检测并移除异常值)。
增加模型的复杂度通常会显著提升模型的方差并减少偏差。反过来,降低模型的复
杂度则会提升模型的偏差并降低方差。这就是为什么称其为权衡。
4.5 正则化线性模型
正如我们在第 1 章和第 2 章中看到的那样,减少过拟合的一个好方法是对模型进行正则
化(即约束模型):它拥有的自由度越少,则过拟合数据的难度就越大。正则化多项式模
型的一种简单方法是减少多项式的次数。
1
对于线性模型,正则化通常是通过约束模型的权重来实现的。现在,我们看一下岭回
归、Lasso 回归和弹性网络,它们实现了三种限制权重的方法。
4.5.1 岭回归
岭回归(也称为 Tikhonov 正则化)是线性回归的正则化版本:将等于
αθ
∑
i=
n
1
i
2
的正则化项
注 8 :不要将这里的偏差概念与线性模型中的偏置项概念弄混。
128
|
第
4
章
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

算法技术手册(原书第2 版)

算法技术手册(原书第2 版)

George T.Heineman, Gary Pollice, Stanley Selkow
管理Kubernetes

管理Kubernetes

Brendan Burns, Craig Tracey

Publisher Resources

ISBN: 9787111665977