book

Practical Synthetic Data Generation

by Khaled El Emam, Lucy Mosquera, Richard Hoptroff

May 2020

Beginner to intermediate

163 pages

4h 31m

English

O'Reilly Media, Inc.

Read now

Unlock full access

Preface
Conventions Used in This BookO’Reilly Online LearningHow to Contact UsAcknowledgments
1. Introducing Synthetic Data Generation
Defining Synthetic DataSynthesis from Real DataSynthesis Without Real DataSynthesis and UtilityThe Benefits of Synthetic DataEfficient Access to DataEnabling Better AnalyticsSynthetic Data as a ProxyLearning to Trust Synthetic DataSynthetic Data Case StudiesManufacturing and DistributionHealthcareFinancial ServicesTransportationSummary
2. Implementing Data Synthesis
When to SynthesizeIdentifiability SpectrumTrade-Offs in Selecting PETs to Enable Data AccessDecision CriteriaPETs ConsideredDecision FrameworkExamples of Applying the Decision FrameworkData Synthesis ProjectsData Synthesis StepsData PreparationThe Data Synthesis PipelineSynthesis Program ManagementSummary
3. Getting Started: Distribution Fitting
Framing DataHow Data Is DistributedFitting Distributions to Real DataGenerating Synthetic Data from a DistributionMeasuring How Well Synthetic Data Fits a DistributionThe Overfitting DilemmaA Little Light WeedingSummary
4. Evaluating Synthetic Data Utility
Synthetic Data Utility Framework: Replication of AnalysisSynthetic Data Utility Framework: Utility MetricsComparing Univariate DistributionsComparing Bivariate StatisticsComparing Multivariate Prediction ModelsDistinguishabilitySummary
5. Methods for Synthesizing Data
Generating Synthetic Data from TheorySampling from a Multivariate Normal DistributionInducing Correlations with Specified Marginal DistributionsCopulas with Known Marginal DistributionsGenerating Realistic Synthetic DataFitting Real Data to Known DistributionsUsing Machine Learning to Fit the DistributionsHybrid Synthetic DataMachine Learning MethodsDeep Learning MethodsSynthesizing SequencesSummary
6. Identity Disclosure in Synthetic Data
Types of DisclosureIdentity DisclosureLearning Something NewAttribute DisclosureInferential DisclosureMeaningful Identity DisclosureDefining Information GainBringing It All TogetherUnique MatchesHow Privacy Law Impacts the Creation and Use of Synthetic DataIssues Under the GDPRIssues Under the CCPAIssues Under HIPAAArticle 29 Working Party OpinionSummary
7. Practical Data Synthesis
Managing Data ComplexityFor Every Pre-Processing Step There Is a Post-Processing StepField TypesThe Need for RulesNot All Fields Have to Be SynthesizedSynthesizing DatesSynthesizing GeographyLookup Fields and TablesMissing Data and Other Data CharacteristicsPartial SynthesisOrganizing Data SynthesisComputing CapacityA Toolbox of TechniquesSynthesizing Cohorts Versus Full DatasetsContinuous Data FeedsPrivacy Assurance as CertificationPerforming Validation Studies to Get Buy-InMotivated Intruder TestsWho Owns Synthetic Data?Conclusions
Index

Content preview from Practical Synthetic Data Generation

Chapter 5. Methods for Synthesizing Data

After describing some basic methods for distribution fitting in the last chapter, we will now use these concepts to generate synthetic data. We will start off with some basic approaches and build up to some more complex ones as the chapter progresses. We will refer to more advanced techniques later on that are beyond the scope of an introductory text, but what we cover should give you a good introduction.

Generating Synthetic Data from Theory

Let’s consider the situation where the analyst does not have any real data to start off with, but has some understanding of the phenomenon that they want to model and generate data for. For example, let’s say that we want to generate data reflecting the relationship between height and weight. It is generally known that height and weight are positively associated.

According to the Centers for Disease Control, the average height for men in the US is approximately 175 cm,¹ and for the sake of our example we will assume a standard deviation of 5 cm. The average weight is 89.7 kg, and we will assume a standard deviation of 10 kg. For the sake of our example, we will model these as normal (Gaussian or bell-shaped) distributions and assume that the correlation between them is 0.5. According to Cohen’s guidelines for the interpretation of effect sizes, a correlation of magnitude equal to 0.5 is considered to be large, 0.3 is considered to be medium, and 0.1 is considered to be small. Any correlation above ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.

Julian F.

Head of Cybersecurity

I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.

Addison B.

Field Engineer

I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.

Amir M.

Data Platform Tech Lead

I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.

Mark W.

Embedded Software Engineer

Publisher Resources

ISBN: 9781492072737Errata Page

Cloud Computing

Data Engineering

Data Science

AI & ML

Programming Languages

Software Architecture

IT/Ops

Security

Design

Business

Soft Skills

Practical Synthetic Data Generation

by Khaled El Emam, Lucy Mosquera, Richard Hoptroff

Chapter 5. Methods for Synthesizing Data

Generating Synthetic Data from Theory

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.