Skip to Content
Spark快速大数据分析(第2版)
book

Spark快速大数据分析(第2版)

by Jules S. Damji, Brooke Wenig, Tathagata Das, Denny Lee
November 2021
Intermediate to advanced
340 pages
10h 46m
Chinese
Posts & Telecom Press
Content preview from Spark快速大数据分析(第2版)
254
10
每个工作节点独立于其他工作节点处理自己的数据分片。如果分区中的数据发生了改变,
那么这个分片(通过
randomSplit()
生成)产生的结果也就不一样了。
既然集群配置和随机种子都是可以固定下来的,我们建议一次性划分数据,然后将数据分
别写入训练集和测试集的文件夹,这样就不用担心不可重现的问题了。
对数据集进行探索性分析时,应该缓存训练集,因为需要在机器学习过程中
多次访问这些数据。具体参见
7.2
节。
10.2.3
 为转化器准备特征
至此,我们已经将数据分为了训练集和测试集。为构建根据卧室数量预测价格的线性
回归模型,现在需要准备数据。在稍后的示例中,我们会引入所有相关特征,目前先
将路走通。线性回归(和
Spark
中的很多其他算法类似)要求所有的输入特征都包含在
DataFrame
的单个向量中。因此,我们需要
转化
数据。
Spark
中的转化器接受
DataFram
e
作为输入,并返回添加了一列或多列的新
DataFrame
。转
化器不从数据中学习,只是用
transform()
方法执行基于规则的转化操作。
对于将所有特征放入单个向量的这个任务,我们使用
VectorAssembler
转化器。
VectorAssembler
接受一串输入列,并创建加了一列的新
DataFrame
,新加的这一列就是特征(
features
)。
这个特征会将所有输入列的值组合到单个向量中。
# Python代码
from pyspark.ml.feature import
VectorAssembler
vecAssembler ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

数据驱动力:企业数据分析实战

数据驱动力:企业数据分析实战

Carl Anderson
数据压缩入门

数据压缩入门

Colt McAnlis, Aleks Haecky
解密金融数据

解密金融数据

Justin Pauley

Publisher Resources

ISBN: 9787115576019