Chapter 6. Identity Disclosure in Synthetic Data
The analysis of privacy risks with synthetic data remains an important topic. In the context of a privacy analysis, we are concerned here with data that pertains to individuals. If the data does not pertain to individuals, then there will be no privacy concerns. For example, if the data pertains to prescriptions or cars, then we would not worry as much about privacy. However, data synthesis is being used extensively to generate data about individuals, and therefore we need to understand the privacy implications.
There is a general belief that synthetic data has negligible privacy risk because there is no unique mapping between the records in the synthetic data and the records in the original data.1 Reiter noted that “identification of units and their sensitive data from synthetic samples is nearly impossible,”2 and Taub et al. said that “it is widely understood that thinking of risk within synthetic data in terms of re-identification, which is how many other SDC [statistical disclosure control] methods approach disclosure risk, is not meaningful.”3
However, in practice, when generating synthetic data it is possible to overfit the synthesis model to the real data, and we have discussed that in earlier chapters of this book. This means that the generated data will look very similar to the original data, hence creating a privacy problem whereby we can map the records in the synthetic data to individuals in the real world. Therefore, ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access