Distribution-Invariant Differential Privacy

Distribution-Invariant Differential Privacy
复制标题

DOI:
10.1016/j.jeconom.2022.05.004
复制
发表时间:
2021-11
影响因子:
6.3
通讯作者:
Xuan Bi;Xiaotong Shen
Xuan Bi;Xiaotong Shen
中科院分区:
经济学2区
文献类型:
--
作者:
Xuan Bi;Xiaotong Shen

文献摘要

相似文献

差异隐私正在成为保护公开共享数据隐私的黄金标准。它被广泛应用于社会科学、数据科学、公共卫生、信息技术和美国十年一次的人口普查。然而,为了保证不同的隐私,现有的方法可能会不可避免地改变原始数据分析的结论,因为私有化往往会改变样本分布。这种现象被称为隐私保护和统计准确性之间的权衡。在这项工作中,我们通过开发一种分布不变私有化(DIP)方法来协调高统计准确性和严格的差分隐私来减轻这种权衡。因此,任何下游统计或机器学习任务都会产生与使用原始数据基本相同的结论。在数字上,在同样严格的隐私保护下,DIP在广泛的模拟研究和真实世界的基准测试中达到了上级的统计准确性。
Differential privacy is becoming one gold standard for protecting the privacy of publicly shared data. It has been widely used in social science, data science, public health, information technology, and the U.S. decennial census. Nevertheless, to guarantee differential privacy, existing methods may unavoidably alter the conclusion of original data analysis, as privatization often changes the sample distribution. This phenomenon is known as the trade-off between privacy protection and statistical accuracy. In this work, we mitigate this trade-off by developing a distribution-invariant privatization (DIP) method to reconcile both high statistical accuracy and strict differential privacy. As a result, any downstream statistical or machine learning task yields essentially the same conclusion as if one used the original data. Numerically, under the same strictness of privacy protection, DIP achieves superior statistical accuracy in in a wide range of simulation studies and real-world benchmarks.