Skip to Content
机器学习速查手册
book

机器学习速查手册

by Matt Harrison
July 2025
Intermediate to advanced
320 pages
3h 10m
Chinese
O'Reilly Media, Inc.
Content preview from 机器学习速查手册

第 9 章 失衡的班级 不平衡的类别

本作品已使用人工智能进行翻译。欢迎您提供反馈和意见:translation-feedback@oreilly.com

如果您正在对数据进行分类,而类的大小并不相对均衡,那么偏向于更受欢迎的类的偏差可能会延续到您的模型中。例如,如果您有 1 个阳性案例和 99 个阴性案例,您只需将所有案例都归类为阴性,就能获得 99% 的准确率。处理不平衡类别有多种方法。

使用不同的衡量标准

一个提示是使用准确率以外的指标(AUC 是一个不错的选择)来校准模型。当目标大小不同时,精确度和召回率也是更好的选择。不过,还可以考虑其他方案。

基于树的算法和集合

基于树的模型可能会有更好的表现,这取决于小类的分布情况。如果它们趋于聚类,就更容易分类。

集合方法可以进一步帮助找出少数类别。随机森林和极梯度提升 (XGBoost) 等树状模型中都有袋式分类和提升方法。

惩罚模型

许多 scikit-learn 分类模型都支持class_weight 参数。将该参数设置为'balanced' 将尝试对少数类进行正则化,激励模型对其进行正确分类。或者,也可以进行网格搜索,并通过传入一个类与权重的字典来指定权重选项(给较小的类更高的权重)。

XGBoost库有一个max_delta_step 参数,可以设置为 1 到 10,使更新步骤更加保守。它还有一个scale_pos_weight 参数,用于设置负样本与正样本的比例(对于二元类)。此外,在分类时,eval_metric 应设置为'auc' ,而不是默认值'error' 。

KNN 模型有一个weights 参数,该参数会使距离较近的邻居产生偏差。如果少数群体样本距离较近,将该参数设置为'distance' 可能会提高性能。

对少数群体进行上采样

您可以通过几种方法对少数类进行增大采样。下面是一个 sklearn 实现:

>>> from sklearn.utils import resample
>>> mask = df.survived == 1
>>> surv_df = df[mask]
>>> death_df = df[~mask]
>>> df_upsample = resample(
...     surv_df,
...     replace=True,
...     n_samples=len(death_df),
...     random_state=42,
... )
>>> df2 = pd.concat([death_df, df_upsample])

>>> df2.survived.value_counts()
1    809
0    809
Name: survived, dtype: int64

我们还可以使用不平衡学习库(imbalanced-learn library)进行随机替换采样:

>>> from imblearn.over_sampling import (
...     RandomOverSampler,
... )
>>> ros = RandomOverSampler(random_state=42)
>>> X_ros, y_ros = ros.fit_sample(X, y)
>>> pd.Series(y_ros).value_counts()
1    809
0    809
dtype: int64

生成少数群体数据

不平衡学习库还可以利用合成少数群体过度采样技术(SMOTE)和自适应合成(ADASYN)采样方法算法生成新的少数群体样本。SMOTE ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

雷达趋势观察:2025年7月

雷达趋势观察:2025年7月

Mike Loukides
R深度学习权威指南

R深度学习权威指南

Posts & Telecom Press, Joshua F. Wiley
Python高级编程(第2版)

Python高级编程(第2版)

Posts & Telecom Press, Michał Jaworski, Tarek Ziadé

Publisher Resources

ISBN: 9798341663046