book

Practical Synthetic Data Generation

by Khaled El Emam, Lucy Mosquera, Richard Hoptroff

May 2020

Beginner to intermediate

163 pages

4h 31m

English

O'Reilly Media, Inc.

Read now

Unlock full access

Preface
Conventions Used in This BookO’Reilly Online LearningHow to Contact UsAcknowledgments
1. Introducing Synthetic Data Generation
Defining Synthetic DataSynthesis from Real DataSynthesis Without Real DataSynthesis and UtilityThe Benefits of Synthetic DataEfficient Access to DataEnabling Better AnalyticsSynthetic Data as a ProxyLearning to Trust Synthetic DataSynthetic Data Case StudiesManufacturing and DistributionHealthcareFinancial ServicesTransportationSummary
2. Implementing Data Synthesis
When to SynthesizeIdentifiability SpectrumTrade-Offs in Selecting PETs to Enable Data AccessDecision CriteriaPETs ConsideredDecision FrameworkExamples of Applying the Decision FrameworkData Synthesis ProjectsData Synthesis StepsData PreparationThe Data Synthesis PipelineSynthesis Program ManagementSummary
3. Getting Started: Distribution Fitting
Framing DataHow Data Is DistributedFitting Distributions to Real DataGenerating Synthetic Data from a DistributionMeasuring How Well Synthetic Data Fits a DistributionThe Overfitting DilemmaA Little Light WeedingSummary
4. Evaluating Synthetic Data Utility
Synthetic Data Utility Framework: Replication of AnalysisSynthetic Data Utility Framework: Utility MetricsComparing Univariate DistributionsComparing Bivariate StatisticsComparing Multivariate Prediction ModelsDistinguishabilitySummary
5. Methods for Synthesizing Data
Generating Synthetic Data from TheorySampling from a Multivariate Normal DistributionInducing Correlations with Specified Marginal DistributionsCopulas with Known Marginal DistributionsGenerating Realistic Synthetic DataFitting Real Data to Known DistributionsUsing Machine Learning to Fit the DistributionsHybrid Synthetic DataMachine Learning MethodsDeep Learning MethodsSynthesizing SequencesSummary
6. Identity Disclosure in Synthetic Data
Types of DisclosureIdentity DisclosureLearning Something NewAttribute DisclosureInferential DisclosureMeaningful Identity DisclosureDefining Information GainBringing It All TogetherUnique MatchesHow Privacy Law Impacts the Creation and Use of Synthetic DataIssues Under the GDPRIssues Under the CCPAIssues Under HIPAAArticle 29 Working Party OpinionSummary
7. Practical Data Synthesis
Managing Data ComplexityFor Every Pre-Processing Step There Is a Post-Processing StepField TypesThe Need for RulesNot All Fields Have to Be SynthesizedSynthesizing DatesSynthesizing GeographyLookup Fields and TablesMissing Data and Other Data CharacteristicsPartial SynthesisOrganizing Data SynthesisComputing CapacityA Toolbox of TechniquesSynthesizing Cohorts Versus Full DatasetsContinuous Data FeedsPrivacy Assurance as CertificationPerforming Validation Studies to Get Buy-InMotivated Intruder TestsWho Owns Synthetic Data?Conclusions
Index

Content preview from Practical Synthetic Data Generation

Chapter 1. Introducing Synthetic Data Generation

We start this chapter by explaining what synthetic data is and its benefits. Artificial intelligence and machine learning (AIML) projects run in various industries, and the use cases that we include in this chapter are intended to give a flavor of the broad applications of data synthesis. We define an AIML project quite broadly as well, to include, for example, the development of software applications that have AIML components.

Defining Synthetic Data

At a conceptual level, synthetic data is not real data, but data that has been generated from real data and that has the same statistical properties as the real data. This means that if an analyst works with a synthetic dataset, they should get analysis results similar to what they would get with real data. The degree to which a synthetic dataset is an accurate proxy for real data is a measure of utility. We refer to the process of generating synthetic data as synthesis.

Data in this context can mean different things. For example, data can be structured data, as one would see in a relational database. Data can also be unstructured text, such as doctors’ notes, transcripts of conversations or online interactions by email or chat. Furthermore, images, videos, audio, and virtual environments are types of data that can be synthesized. Using machine learning, it is possible to create realistic pictures of people who do not exist in the real world.

There are three types of synthetic data. ...

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

O’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.

Julian F.

Head of Cybersecurity

I wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.

Addison B.

Field Engineer

I’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.

Amir M.

Data Platform Tech Lead

I'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.

Mark W.

Embedded Software Engineer

Publisher Resources

ISBN: 9781492072737Errata Page

Cloud Computing

Data Engineering

Data Science

AI & ML

Programming Languages

Software Architecture

IT/Ops

Security

Design

Business

Soft Skills

Practical Synthetic Data Generation

by Khaled El Emam, Lucy Mosquera, Richard Hoptroff

Chapter 1. Introducing Synthetic Data Generation

Defining Synthetic Data

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.