Chapter 4. The Future of Data Synthesis
While significant progress has been made over the last few years in making synthetic data generation practical and scalable, we need some additional requirements for future work and improvements to the current state of practice. This chapter is a summary of the key issues that need to be worked on. It does not present a research and development agenda, but rather a set of items to consider when developing such an agenda.
We cover four main issues. First, we need to develop a data utility framework. Such a framework would make it easier to benchmark various data synthesis techniques. The second issue, which is coming up more frequently, is the need to remove certain relationships from synthetic data for commercial or security reasons. Third, data watermarking will become increasingly important as more synthetic data is generated and shared. Finally, simulators that can generate different types of synthetic data would provide powerful capabilities.
Creating a Data Utility Framework
As discussed in Chapter 1, data utility is important for the adoption of synthetic data. The higher the data utility of synthetic data, the greater the number of use cases where it would be a good tool to accelerate AIML efforts, and the more likely that analysts will be comfortable using it.
In practice, we are seeing that a significant dataset, the 2020 decennial US census, is being shared as synthetic data and derivatives from synthetic data. The question of ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access