Skip to Content
Spark快速大数据分析(第2版)
book

Spark快速大数据分析(第2版)

by Jules S. Damji, Brooke Wenig, Tathagata Das, Denny Lee
November 2021
Intermediate to advanced
340 pages
10h 46m
Chinese
Posts & Telecom Press
Content preview from Spark快速大数据分析(第2版)
MLlib
实现机器学习
251
10.2
 设计机器学习流水线
本节将介绍如何创建和调优机器学习流水线。很多机器学习框架有流水线的概念,其表示
对数据组织一系列操作的一种方式。
MLlib
Pipeline
API
提供了一套构建在
DataFrame
上的高层
API
,以便用户组织机器学习工作流
Pipeline API
是由一系列转化器和预估器组
成的,稍后我们会深入讨论。
本章会使用来自
Inside Airbnb
的旧金山住房数据集
。这个数据集包含旧金山的
Airbnb
租信息,比如卧室数量、位置、评分,等等。我们的目标是构建出预测这个城市的挂牌房
源每晚出租价格的模型。这是一个回归问题,因为价格是连续的变量。我们将体验一个数
据科学家解决这类问题所需经历的工作流程,包括特征工程、模型构建、超参数调优,以
及模型效果评估。这个数据集比较乱,建模难度较大。(现实世界中的很多数据集是这样
的!)因此,如果你自己尝试建模,请不要因为早期模型效果不佳而沮丧。
本章的目的不是展示
MLlib
的各个
API
,而是帮助你获得使用
MLlib
所需要的技能和知
识,以便构建端到端的流水线。在深入探讨细节前,我们先学习
MLlib
的一些术语定义。
转化器
它接受一个
DataFrame
为输入,并返回添加了一列或多列的新
DataFrame
。转化器不
会从数据中学到任何参数,只是简单地应用基于规则的转化操作,从而为模型训练准备
数据,或使用训练好的
MLlib
模型来生成预测结果。转化器有
.transform()
方法。
预估器
DataFrame ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

数据驱动力:企业数据分析实战

数据驱动力:企业数据分析实战

Carl Anderson
数据压缩入门

数据压缩入门

Colt McAnlis, Aleks Haecky
解密金融数据

解密金融数据

Justin Pauley

Publisher Resources

ISBN: 9787115576019