Privacy of Synthetic Data: A Statistical Framework

Privacy of Synthetic Data: A Statistical Framework
复制标题

DOI:
10.1109/tit.2022.3216793
复制
发表时间:
2021-09
影响因子:
2.5
通讯作者:
M. Boedihardjo;T. Strohmer;R. Vershynin
M. Boedihardjo;T. Strohmer;R. Vershynin
中科院分区:
计算机科学2区
文献类型:
--
作者:
M. Boedihardjo;T. Strohmer;R. Vershynin

文献摘要

相似文献

隐私保护数据分析正成为一个具有深远影响的挑战性问题。特别是,合成数据是解决数据隐私和数据共享之间矛盾困境的一个有前景的概念。然而,众所周知,准确生成某些类型的隐私合成数据是NP难问题。我们为差分隐私合成数据开发了一个统计框架,这使我们能够规避该问题的计算难度。我们将真实数据视为根据某种未知密度从总体$\Omega$中抽取的随机样本。然后,我们用一个小得多的随机子集$\Omega^{\ast}$来代替$\Omega$,我们根据某种已知密度对其进行抽样。我们通过拟合从真实数据中获得的特定线性统计量,在简化空间$\Omega^{\ast}$上生成合成数据。为确保隐私,我们使用常见的拉普拉斯机制。利用雷尼条件数的概念(它衡量抽样分布与总体分布的相关性程度),我们推导出了所提方法提供的隐私和准确性的明确界限。
Privacy-preserving data analysis is emerging as a challenging problem with far-reaching impact. In particular, synthetic data are a promising concept toward solving the aporetic conflict between data privacy and data sharing. Yet, it is known that accurately generating private, synthetic data of certain kinds is NP-hard. We develop a statistical framework for differentially private synthetic data, which enables us to circumvent the computational hardness of the problem. We consider the true data as a random sample drawn from a population $\Omega $ according to some unknown density. We then replace $\Omega $ by a much smaller random subset $\Omega ^{\ast}$ , which we sample according to some known density. We generate synthetic data on the reduced space $\Omega ^{\ast}$ by fitting the specified linear statistics obtained from the true data. To ensure privacy we use the common Laplacian mechanism. Employing the concept of Rényi condition number, which measures how well the sampling distribution is correlated with the population distribution, we derive explicit bounds on the privacy and accuracy provided by the proposed method.