第 11 章 模型选择 模型选择
本作品已使用人工智能进行翻译。欢迎您提供反馈和意见:translation-feedback@oreilly.com
本章将讨论优化超参数。本章还将探讨模型是否需要更多数据才能表现更好的问题。
验证曲线
创建验证曲线是确定超参数适当值的一种方法。 验证曲线是显示模型性能如何响应超参数值变化的曲线图(见图 11-1)。图表同时显示了训练数据和验证数据。通过验证分数,我们可以推断出模型对未知数据的响应情况。通常情况下,我们会选择一个能使验证得分最大化的超参数。
在下面的示例中,我们将使用 Yellowbrick 来观察改变max_depth超级参数的值是否会改变随机森林的模型性能。你可以提供一个scoring 参数设置为 scikit-learn 模型度量(分类的默认值是'accuracy' ):
提示
使用n_jobs 参数可充分利用 CPU,加快运行速度。如果将其设置为-1 ,则会使用所有 CPU。
>>>fromyellowbrick.model_selectionimport(...ValidationCurve,...)>>>fig,ax=plt.subplots(figsize=(6,4))>>>vc_viz=ValidationCurve(...RandomForestClassifier(n_estimators=100),...param_name="max_depth",...param_range=np.arange(1,11),...cv=10,...n_jobs=-1,...)>>>vc_viz.fit(X,y)>>>vc_viz.poof()>>>fig.savefig("images/mlpr_1101.png",dpi=300)
图 11-1. 验证曲线报告。
ValidationCurve 类支持一个scoring 参数。该参数可以是一个自定义函数,也可以是以下选项之一,具体取决于任务。
分类scoring 选项包括'accuracy','average_precision','f1','f1_micro','f1_macro','f1_weighted','f1_samples','neg_log_loss','precision','recall', 和'roc_auc' 。
聚类scoring 选项:'adjusted_mutual_info_score','adjusted_rand_score','completeness_score' 、'fowlkesmallows_score','homogeneity_score','mutual_info_score','normalized_mutual_info_score', 和'v_measure_score' 。
回归scoring 选项 : 'explained_variance','neg_mean_absolute_error','neg_mean_squared_error','neg_mean_squared_log_error','neg_median_absolute_error', 和'r2' 。
学习曲线
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access