book

精通特征工程

Name: 精通特征工程
ISBN: 9787115509680

by Alice Zheng, Amanda Casari

April 2019

Intermediate to advanced

172 pages

4h 39m

Chinese

Posts & Telecom Press

Read now

Unlock full access

致谢
Alice的特别感谢
Amanda的特别感谢
电子书
第1章机器学习流程
1.1 数据1.2 任务
1.3 模型
1.4 特征
1.5 模型评价
第2章简单而又奇妙的数值
2.1 标量、向量和空间
2.2 处理计数
2.2.1 二值化
2.2.2 区间量化（分箱）
2.3 对数变换
2.3.1 对数变换实战
2.3.2 指数变换：对数变换的推广
2.4 特征缩放/归一化
2.4.1 min-max缩放2.4.2 特征标准化/方差缩放
2.4.3 ℓ 2归一化
2.5 交互特征
2.6 特征选择
2.7 小结
2.8 参考文献
第3章文本数据：扁平化、过滤和分块
3.1 元素袋：将自然文本转换为扁平向量
3.1.1 词袋
3.1.2 n元词袋
3.2 使用过滤获取清洁特征
3.2.1 停用词
3.2.2 基于频率的过滤
3.2.3 词干提取
3.3 意义的单位：从单词、n元词到短语
3.3.1 解析与分词
3.3.2 通过搭配提取进行短语检测 (1/2)
3.3.2 通过搭配提取进行短语检测 (2/2)
3.4 小结
3.5 参考文献
第4章特征缩放的效果：从词袋到tf-idf
4.1 tf-idf：词袋的一种简单扩展
4.2 tf-idf方法测试
4.2.1 创建分类数据集
4.2.2 使用tf-idf变换来缩放词袋
4.2.3 使用逻辑回归进行分类
4.2.4 使用正则化对逻辑回归进行调优
4.3 深入研究：发生了什么
4.4 小结
4.5 参考文献
第5章分类变量：自动化时代的数据计数
5.1 分类变量的编码
5.1.1 one-hot编码5.1.2 虚拟编码
5.1.3 效果编码
5.1.4 各种分类变量编码的优缺点
5.2 处理大型分类变量
5.2.1 特征散列化
5.2.2 分箱计数 (1/2)
5.2.2 分箱计数 (2/2)
5.3 小结
5.4 参考文献
第6章数据降维：使用PCA挤压数据
6.1 直观理解
6.2 数学推导
6.2.1 线性投影
6.2.2 方差和经验方差
6.2.3 主成分：第一种表示形式
6.2.4 主成分：矩阵-向量表示形式6.2.5 主成分的通用解
6.2.6 特征转换
6.2.7 PCA实现
6.3 PCA实战
6.4 白化与ZCA
6.5 PCA的局限性与注意事项
6.6 用例
6.7 小结
6.8 参考文献
第7章非线性特征化与k-均值模型堆叠
7.1 k-均值聚类
7.2 使用聚类进行曲面拼接
7.3 用于分类问题的k-均值特征化
7.4 优点、缺点以及陷阱
7.5 小结
7.6 参考文献
第8章自动特征生成：图像特征提取和深度学习
8.1 最简单的图像特征（以及它们因何失效）
8.2 人工特征提取：SIFT和HOG
8.2.1 图像梯度
8.2.2 梯度方向直方图
8.2.3 SIFT体系
8.3 通过深度神经网络学习图像特征
8.3.1 全连接层
8.3.2 卷积层
8.3.3 ReLU变换
8.3.4 响应归一化层
8.3.5 池化层
8.3.6 AlexNet的结构
8.4 小结
8.5 参考文献
第9章回到特征：建立学术论文推荐器
9.1 基于项目的协同过滤
9.2 第一关：数据导入、清理和特征解析 (1/2)
9.2 第一关：数据导入、清理和特征解析 (2/2)
9.3 第二关：更多特征工程和更智能的模型
9.4 第三关：更多特征=更多信息
9.5 小结
9.6 参考文献
附录A 线性建模与线性代数基础
A.1 线性分类概述
A.2 矩阵的解析
A.2.1 从向量到子空间
A.2.2 奇异值分解（SVD）
A.2.3 数据矩阵的四个基本子空间
A.3 线性系统求解
A.4 参考文献
作者简介
封面简介

Content preview from 精通特征工程

｜

第

章

Test score with bow features: 0.775873066497

Test score with l2-normalized features: 0.763514590974

Test score with tf-idf features: 0.743182905438

事与愿违，结果显示准确率最高的分类器使用的是词袋特征，这真是出乎意料。实际上，

出现这种情况的原因在于分类器没有很好地“调优”，这是在比较分类器时经常犯的错误。

4.2.4

　使用正则化对逻辑回归进行调优

逻辑回归有些不切实际的功能。当特征数量大于数据点数量时，找出最佳模型这个问题就

变得

不确定

了。解决这个问题的一种方法是在训练过程中加入额外的限制条件，这就是

正

则化

，本节将讨论它的技术细节。

逻辑回归的大多数具体实现都允许正则化。要使用正则化，必须确定一个正则化参数。正

则化参数是一种

超参数

，不能在模型训练过程中自动学习。相反，它们必须根据具体的问

题进行调优，并提供给训练算法，这个过程就是超参数调优。（如果想详细了解如何评价

机器学习模型，参见

Zheng (2015)

。）一种基本的超参数调优方法称为

网格搜索

：先确定一

个超参数网格，然后使用调优程序自动搜索，找到网格中的最优超参数设置。找到最优超

参数设置之后，你可以使用该设置在整个训练集上训练一个模型，然后使用它在测试集上

的表现作为这类模型的最终评价。

重要

：

比较模型时要对超参数进行调优

当比较模型或特征时，必须对超参数进行调优。软件包的默认设置总是能返 ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.

Julian F.

Head of Cybersecurity

I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.

Addison B.

Field Engineer

I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.

Amir M.

Data Platform Tech Lead

I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.

Mark W.

Embedded Software Engineer

Publisher Resources

ISBN: 9787115509680

Cloud Computing

Data Engineering

Data Science

AI & ML

Programming Languages

Software Architecture

IT/Ops

Security

Design

Business

Soft Skills

精通特征工程

by Alice Zheng, Amanda Casari

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.