book

Principles of Data Wrangling

by Joseph M. Hellerstein, Tye Rattenbury, Jeffrey Heer, Sean Kandel, Connor Carreras

July 2017

Beginner

92 pages

2h 29m

English

O'Reilly Media, Inc.

Read now

Unlock full access

Foreword
1. Introduction
Magic Thresholds, PYMK, and User Growth at Facebook
2. A Data Workflow Framework
How Data Flows During and Across ProjectsConnecting Analytic Actions to Data Movement: A Holistic Workflow Framework for Data ProjectsRaw Data Stage Actions: Ingest Data and Create MetadataIngesting Known and Unknown DataCreating MetadataRefined Data Stage Actions: Create Canonical Data and Conduct Ad Hoc AnalysesDesigning Refined DataRefined Stage Analytical ActionsProduction Data Stage Actions: Create Production Data and Build Automated SystemsCreating Optimized DataDesigning Regular Reports and Automated Products/ServicesData Wrangling within the Workflow Framework
3. The Dynamics of Data Wrangling
Data Wrangling DynamicsAdditional Aspects: Subsetting and SamplingCore Transformation and Profiling ActionsData Wrangling in the Workflow FrameworkIngesting DataDescribing DataAssessing Data UtilityDesigning and Building Refined DataAd Hoc ReportingExploratory Modeling and ForecastingBuilding an Optimized DatasetRegular Reporting and Building Data-Driven Products and Services
4. Profiling
Overview of ProfilingIndividual Value Profiling: Syntactic ProfilingIndividual Value Profiling: Semantic ProfilingSet-Based ProfilingProfiling Individual Values in the Candidate Master FileSyntactic Profiling in the Candidate Master FileSet-Based Profiling in the Candidate Master File
5. Transformation: Structuring
Overview of StructuringIntrarecord Structuring: Extracting ValuesPositional ExtractionPattern ExtractionComplex Structure ExtractionIntrarecord Structuring: Combining Multiple Record FieldsInterrecord Structuring: Filtering Records and FieldsInterrecord Structuring: Aggregations and PivotsSimple AggregationsColumn-to-Row PivotsRow-to-Column Pivots
6. Transformation: Enriching
UnionsJoinsInserting MetadataDerivation of ValuesGenericProprietary
7. Using Transformation to Clean Data
Addressing Missing/NULL ValuesAddressing Invalid Values
8. Roles and Responsibilities
Skills and ResponsibilitiesData EngineerData ArchitectData ScientistAnalystRoles Across the Data Workflow FrameworkOrganizational Best Practices
9. Data Wrangling Tools
Data Size and InfrastructureData StructuresExcelSQLTrifacta WranglerTransformation ParadigmsExcelSQLTrifacta WranglerChoosing a Data Wrangling Tool

Content preview from Principles of Data Wrangling

Chapter 1. Introduction

Let’s begin with the most important question: why should you read this book? The answer is simple: you want more value from your data. To put a little more meat on that statement, our objective in writing this book is to help the variety of people who manage the analysis or application of data in their organizations. The data might or might not be “yours,” in the strict sense of ownership. But the pains in extracting value from this data are.

We’re focused on two kinds of readers. First are people who manage the analysis and application of data indirectly—the managers of teams or directors of data projects. Second are people who work with data directly—the analysts, engineers, architects, statisticians, and scientists.

If you’re reading this book, you’re interested in extracting value from data. We can categorize this value into two types along a temporal dimension: near-term value and long-term value. In the near term, you likely have a sizable list of questions that you want to answer using your data. Some of these questions might be vague; for example, “Are people really shifting toward interacting with us through their mobile devices?” Other questions might be more specific: “When will our customers’ interactions primarily originate from mobile devices instead of from desktops or laptops?”

What is stopping you from answering these questions? The most common answer we hear is “time.” You know the questions, you know how to answer them, but you just don’t ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.

Julian F.

Head of Cybersecurity

I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.

Addison B.

Field Engineer

I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.

Amir M.

Data Platform Tech Lead

I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.

Mark W.

Embedded Software Engineer

Publisher Resources

ISBN: 9781491938911Errata Page

Cloud Computing

Data Engineering

Data Science

AI & ML

Programming Languages

Software Architecture

IT/Ops

Security

Design

Business

Soft Skills

Principles of Data Wrangling

by Joseph M. Hellerstein, Tye Rattenbury, Jeffrey Heer, Sean Kandel, Connor Carreras

Chapter 1. Introduction

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.