Comparative Study of Differentially Private Data Synthesis Methods

Comparative Study of Differentially Private Data Synthesis Methods
复制标题

DOI:
10.1214/19-sts742
复制
发表时间:
2020-05-01
影响因子:
5.7
通讯作者:
Liu, Fang
Liu, Fang
中科院分区:
数学2区
文献类型:
--
作者:
Bowen, Claire McKay;Liu, Fang

文献摘要

被引文献

相似文献

在研究人员之间共享数据或发布数据供公众使用时,存在暴露数据集中个人敏感信息的风险。数据合成是一种统计披露限制技术,用于发布具有伪个体记录的合成数据集。传统的数据综合技术通常依赖于对数据入侵者的行为和背景知识的强假设来评估披露风险。差分隐私(DP)制定了一个强大的和鲁棒的隐私保证在数据发布的理论方法,而不必建模入侵者的行为。已作出努力,将发展伙伴关系概念纳入数据综合进程。在本文中,我们研究目前的差异私有数据合成(DIPS)技术,释放个人层面的替代数据的原始数据,比较技术的概念和评估的统计效用和推理性能的合成数据,通过每个DIPS技术通过广泛的模拟研究。我们的工作揭示了各种DIPS方法的实际可行性和实用性,并提出了未来的研究方向。
When sharing data among researchers or releasing data for public use, there is a risk of exposing sensitive information of individuals in the data set. Data synthesis is a statistical disclosure limitation technique for releasing synthetic data sets with pseudo individual records. Traditional data synthesis techniques often rely on strong assumptions of a data intruder's behaviors and background knowledge to assess disclosure risk. Differential privacy (DP) formulates a theoretical approach for a strong and robust privacy guarantee in data release without having to model intruders' behaviors. Efforts have been made aiming to incorporate the DP concept in the data synthesis process. In this paper, we examine current DIfferentially Private Data Synthesis (DIPS) techniques for releasing individual-level surrogate data for the original data, compare the techniques conceptually and evaluate the statistical utility and inferential properties of the synthetic data via each DIPS technique through extensive simulation studies. Our work sheds light on the practical feasibility and utility of the various DIPS approaches, and suggests future research directions for DIPS.