计算机科学
差别隐私
任务(项目管理)
数据挖掘
合成数据
领域(数学分析)
样品(材料)
数据建模
人工智能
机器学习
数据库
经济
数学
化学
色谱法
数学分析
管理
作者
Zhikun Zhang,Tianhao Wang,Ninghui Li,Jean Honorio,Michael Backes,Shibo He,Jiming Chen,Yang Zhang
标识
DOI:10.48550/arxiv.2012.15128
摘要
In differential privacy (DP), a challenging problem is to generate synthetic datasets that efficiently capture the useful information in the private data. The synthetic dataset enables any task to be done without privacy concern and modification to existing algorithms. In this paper, we present PrivSyn, the first automatic synthetic data generation method that can handle general tabular datasets (with 100 attributes and domain size $>2^{500}$). PrivSyn is composed of a new method to automatically and privately identify correlations in the data, and a novel method to generate sample data from a dense graphic model. We extensively evaluate different methods on multiple datasets to demonstrate the performance of our method.
科研通智能强力驱动
Strongly Powered by AbleSci AI