book

Bad Data Handbook

Name: Bad Data Handbook
Author: Q. Ethan McCallum
ISBN: 9781449324971

by Q. Ethan McCallum

November 2012

Beginner to intermediate

264 pages

7h 48m

English

O'Reilly Media, Inc.

Read now

Unlock full access

Bad Data Handbook
SPECIAL OFFER: Upgrade this ebook with O’Reilly
About the Authors
Preface
Conventions Used in This BookUsing Code ExamplesSafari® Books OnlineHow to Contact UsAcknowledgments
1. Setting the Pace: What Is Bad Data?
2. Is It Just Me, or Does This Data Smell Funny?
Understand the Data StructureField ValidationValue ValidationPhysical Interpretation of Simple StatisticsVisualizationKeyword PPC ExampleSearch Referral ExampleRecommendation AnalysisTime Series DataConclusion
3. Data Intended for Human Consumption, Not Machine Consumption
The DataThe Problem: Data Formatted for Human ConsumptionThe Arrangement of DataData Spread Across Multiple FilesThe Solution: Writing CodeReading Data from an Awkward FormatReading Data Spread Across Several FilesPostscriptOther FormatsSummary
4. Bad Data Lurking in Plain Text
Which Plain Text Encoding?Guessing Text EncodingNormalizing TextProblem: Application-Specific Characters Leaking into Plain TextText Processing with PythonExercises
5. (Re)Organizing the Web’s Data
Can You Get That?General Workflow Examplerobots.txtIdentifying the Data Organization PatternStore Offline Version for ParsingScrape the Information Off the PageThe Real DifficultiesDownload the Raw Content If PossibleForms, Dialog Boxes, and New WindowsFlashThe Dark SideConclusion
6. Detecting Liars and the Confused in Contradictory Online Reviews
WeottaGetting ReviewsSentiment ClassificationPolarized LanguageCorpus CreationTraining a ClassifierValidating the ClassifierDesigning with DataLessons LearnedSummaryResources

7. Will the Bad Data Please Stand Up?
Example 1: Defect Reduction in ManufacturingExample 2: Who’s Calling?Example 3: When “Typical” Does Not Mean “Average”Lessons LearnedWill This Be on the Test?
8. Blood, Sweat, and Urine
A Very Nerdy Body Swap ComedyHow Chemists Make Up NumbersAll Your Database Are Belong to UsCheck, PleaseLive Fast, Die Young, and Leave a Good-Looking Corpse Code RepositoryRehab for Chemists (and Other Spreadsheet Abusers)tl;dr
9. When Data and Reality Don’t Match
Whose Ticker Is It Anyway?Splits, Dividends, and RescalingBad RealityConclusion
10. Subtle Sources of Bias and Error
Imputation Bias: General IssuesReporting Errors: General IssuesOther Sources of BiasTopcoding/BottomcodingSeam BiasProxy ReportingSample SelectionConclusionsReferences
11. Don’t Let the Perfect Be the Enemy of the Good: Is Bad Data Really Bad?
But First, Let’s Reflect on Graduate School …Moving On to the Professional WorldMoving into Government WorkGovernment Data Is Very RealService Call Data as an Applied ExampleMoving ForwardLessons Learned and Looking Ahead
12. When Databases Attack: A Guide for When to Stick to Files
HistoryBuilding My ToolsetThe Roadblock: My DatastoreConsider Files as Your DatastoreFiles Are Simple!Files Work with EverythingFiles Can Contain Any Data TypeData Corruption Is LocalThey Have Great ToolingThere’s No Install TaxFile ConceptsEncodingText FilesBinary DataMemory-Mapped FilesFile FormatsDelimitersA Web Framework Backed by FilesMotivationImplementationReflections
13. Crouching Table, Hidden Network
A Relational Cost Allocations ModelThe Delicate Sound of a Combinatorial Explosion…The Hidden Network EmergesStoring the GraphNavigating the Graph with GremlinFinding Value in Network PropertiesThink in Terms of Multiple Data Models and Use the Right Tool for the JobAcknowledgments
14. Myths of Cloud Computing
Introduction to the CloudWhat Is “The Cloud”?The Cloud and Big DataIntroducing FredAt First Everything Is GreatThey Put 100% of Their Infrastructure in the CloudAs Things Grow, They Scale Easily at FirstThen Things Start Having TroubleThey Need to Improve PerformanceHigher IO Becomes CriticalA Major Regional Outage Causes Massive DowntimeHigher IO Comes with a CostData Sizes IncreaseGeo Redundancy Becomes a PriorityHorizontal Scale Isn’t as Easy as They HopedCosts Increase DramaticallyFred’s FolliesMyth 1: Cloud Is a Great Solution for All Infrastructure ComponentsHow This Myth Relates to Fred’s StoryMyth 2: Cloud Will Save Us MoneyHow This Myth Relates to Fred’s StoryMyth 3: Cloud IO Performance Can Be Improved to Acceptable Levels Through Software RAIDHow This Myth Relates to Fred’s StoryMyth 4: Cloud Computing Makes Horizontal Scaling EasyHow This Myth Relates to Fred’s StoryConclusion and Recommendations
15. The Dark Side of Data Science
Avoid These PitfallsKnow Nothing About Thy DataBe Inconsistent in Cleaning and Organizing the DataAssume Data Is Correct and CompleteSpillover of Time-Bound DataThou Shalt Provide Your Data Scientists with a Single Tool for All TasksUsing a Production Environment for Ad-Hoc AnalysisThe Ideal Data Science EnvironmentThou Shalt Analyze for Analysis’ Sake OnlyThou Shalt Compartmentalize LearningsThou Shalt Expect Omnipotence from Data ScientistsWhere Do Data Scientists Live Within the Organization?Final Thoughts
16. How to Feed and Care for Your Machine-Learning Experts
Define the ProblemFake It Before You Make ItCreate a Training SetPick the FeaturesEncode the DataSplit Into Training, Test, and Solution SetsDescribe the ProblemRespond to QuestionsIntegrate the SolutionsConclusion
17. Data Traceability
Why?Personal ExperienceSnapshottingSaving the SourceWeighting SourcesBacking Out DataSeparating Phases (and Keeping them Pure)Identifying the Root CauseFinding Areas for ImprovementImmutability: Borrowing an Idea from Functional ProgrammingAn ExampleCrawlersChangeClusteringPopularityConclusion
18. Social Media: Erasable Ink?
Social Media: Whose Data Is This Anyway?ControlCommercial ResyndicationExpectations Around Communication and ExpressionTechnical Implications of New End User ExpectationsWhat Does the Industry Do?Validation APIUpdate Notification APIWhat Should End Users Do?How Do We Work Together?
19. Data Quality Analysis Demystified: Knowing When Your Data Is Good Enough
Framework Introduction: The Four Cs of Data Quality AnalysisCompleteCoherentCorrectaCcountableConclusion
Index
About the Author
Colophon
SPECIAL OFFER: Upgrade this ebook with O’Reilly
Copyright

Content preview from Bad Data Handbook

Chapter 1. Setting the Pace: What Is Bad Data?

We all say we like data, but we don’t.

We like getting insight out of data. That’s not quite the same as liking the data itself.

In fact, I dare say that I don’t quite care for data. It sounds like I’m not alone.

It’s tough to nail down a precise definition of “Bad Data.” Some people consider it a purely hands-on, technical phenomenon: missing values, malformed records, and cranky file formats. Sure, that’s part of the picture, but Bad Data is so much more. It includes data that eats up your time, causes you to stay late at the office, drives you to tear out your hair in frustration. It’s data that you can’t access, data that you had and then lost, data that’s not the same today as it was yesterday…

In short, Bad Data is data that gets in the way. There are so many ways to get there, from cranky storage, to poor representation, to misguided policy. If you stick with this data science bit long enough, you’ll certainly encounter your fair share.

To that end, we decided to compile Bad Data Handbook, a rogues gallery of data troublemakers. We found 19 people from all reaches of the data arena to talk about how data issues have bitten them, and how they’ve healed.

In particular:

Guidance for Grubby, Hands-on Work

You can’t assume that a new dataset is clean and ready for analysis. Kevin Fink’s Is It Just Me, or Does This Data Smell Funny? (Chapter 2) offers several techniques to take the data for a test drive.

There’s plenty of data trapped in spreadsheets, ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.

Julian F.

Head of Cybersecurity

I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.

Addison B.

Field Engineer

I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.

Amir M.

Data Platform Tech Lead

I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.

Mark W.

Embedded Software Engineer

Publisher Resources

ISBN: 9781449324957Errata

Cloud Computing

Data Engineering

Data Science

AI & ML

Programming Languages

Software Architecture

IT/Ops

Security

Design

Business

Soft Skills

Bad Data Handbook

by Q. Ethan McCallum

Chapter 1. Setting the Pace: What Is Bad Data?

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.