Skip to Content
Beautiful Data
book

Beautiful Data

by Toby Segaran, Jeff Hammerbacher
July 2009
Beginner to intermediate
384 pages
12h 56m
English
O'Reilly Media, Inc.
Content preview from Beautiful Data

Preprocessing the Data

We'll start from the beginning: like many websites, FaceStat runs on an SQL database. The judgment interface takes user judgments and saves them as a set of (face ID, attribute, judgment) triples. The first thing we do is extract those 10 million rows from the database. This gives us a file that looks like:

face_id   key          value
149777    describe     serious
18717     trustworthy  3
140467    attractive   2
149777    describe     five-head
...

We're interested in exploring the relationships between different types of perceived attributes. One interesting question is, "How old do I look?" The very first thing to do is to look at the responses that people have given. Unix command-line tools make it easy to quickly see a histogram of responses. The most common responses look like reasonable ages, but we also see a problem:

Look at                    $ cat data.tsv  |
age judgments'                 grep "age"  |
values                         cut -f3  |
and count how many times       sort  |
each value occurs,             uniq -c  |
and order by this count.       sort -nr

Here's the output of this shell pipeline. For each line, the first number is the frequency count. The second string is the response value—exactly what the user typed in the web form in response to the question How old do I look? Most often, she typed in a number, but there are some issues:

70472 19
70021 22
69387 18
68423 17
...
27 24\r\n
27 17\r\n
23 01
21 16\r\n
...
1 old enough to know better
1 hopefully over 21
1 e
1 ??
...

FaceStat has existed for eight months and undergone many changes, so data has been collected ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

Praktische Statistik für Data Scientists, 2nd Edition

Praktische Statistik für Data Scientists, 2nd Edition

Peter Bruce, Andrew Bruce, Peter Gedeck
Beautiful Visualization

Beautiful Visualization

Julie Steele, Noah Iliinsky
Werde ein Data Head

Werde ein Data Head

Alex J. Gutman, Jordan Goldmeier
Basiswissen für Softwarearchitekten, 4th Edition

Basiswissen für Softwarearchitekten, 4th Edition

Mahbouba Gharbi, Arne Koschel, Andreas Rausch, Gernot Starke

Publisher Resources

ISBN: 9780596801656Catalog PageErrata