Private sampling: a noiseless approach for generating differentially private synthetic data

Private sampling: a noiseless approach for generating differentially private synthetic data
复制标题

DOI:
10.1137/21m1449944
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
M. Boedihardjo;T. Strohmer;R. Vershynin
M. Boedihardjo;T. Strohmer;R. Vershynin
中科院分区:
其他
文献类型:
--
作者:
M. Boedihardjo;T. Strohmer;R. Vershynin

文献摘要

相似文献

在一个人工智能和数据科学变得无处不在的世界里,数据共享越来越多地与数据隐私问题发生冲突。差别隐私已成为一个严格的框架,用于保护统计数据库中的个人隐私,同时发布有关数据库的有用统计信息。实现差异隐私的标准方法是向数据中注入足够数量的噪声。然而,除了差异隐私的其他限制外,这种添加噪声的过程将影响数据的准确性和实用性。另一种在数据共享中实现隐私的方法是基于合成数据的概念。合成数据的目标是创建一个尽可能真实的数据集,不仅保持原始数据的细微差别,而且这样做时没有暴露敏感信息的风险。将差异隐私与合成数据相结合被认为是两全其美的解决方案。在这项工作中,我们提出了第一种构造差分私有合成数据的去噪方法;我们通过一种称为“私有采样”的机制来实现这一点。使用布尔立方体作为基准数据模型,我们得到了所构造的合成数据的精度和保密性的显式界。关键的数学工具是超收缩、二元性和经验过程。我们的私人抽样机制的一个核心要素是严格的“边际修正”方法,它具有一个显著的特性,即可以利用重要性重新加权来精确匹配样本的边际和总体的边际。
In a world where artificial intelligence and data science become omnipresent, data sharing is increasingly locking horns with data-privacy concerns. Differential privacy has emerged as a rigorous framework for protecting individual privacy in a statistical database, while releasing useful statistical information about the database. The standard way to implement differential privacy is to inject a sufficient amount of noise into the data. However, in addition to other limitations of differential privacy, this process of adding noise will affect data accuracy and utility. Another approach to enable privacy in data sharing is based on the concept of synthetic data. The goal of synthetic data is to create an as-realistic-as-possible dataset, one that not only maintains the nuances of the original data, but does so without risk of exposing sensitive information. The combination of differential privacy with synthetic data has been suggested as a best-of-both-worlds solutions. In this work, we propose the first noisefree method to construct differentially private synthetic data; we do this through a mechanism called"private sampling". Using the Boolean cube as benchmark data model, we derive explicit bounds on accuracy and privacy of the constructed synthetic data. The key mathematical tools are hypercontractivity, duality, and empirical processes. A core ingredient of our private sampling mechanism is a rigorous"marginal correction"method, which has the remarkable property that importance reweighting can be utilized to exactly match the marginals of the sample to the marginals of the population.