Differentially private data release via statistical election to partition sequentially Statistical election to partition sequentially

Differentially private data release via statistical election to partition sequentially Statistical election to partition sequentially
复制标题

DOI:
10.1007/s40300-021-00201-0
复制
发表时间:
2021-03-12
影响因子:
0.8
通讯作者:
Su, Bingyue
Su, Bingyue
中科院分区:
其他
文献类型:
--
作者:
Bowen, Claire McKay;Liu, Fang;Su, Bingyue

文献摘要

被引文献

相似文献

差分隐私(DP)将隐私以数学形式形式化,并为隐私保护提供了一个健壮的概念。差分私有数据合成(DIPS)技术在DP框架中生成和发布合成的个人级数据。开发DIPS方法的一个关键挑战是保存合成数据的统计效用,特别是在高维环境中。我们提出了一种新的DIPS方法,统计选择分区顺序(STEPS),该方法根据实际或统计重要性度量根据属性的重要性排序对数据进行分区。STEPS的目的是对重要等级较高的属性实现更好的原始信息保存,从而产生更有用的综合数据。我们提出了一种实现STEPS过程的算法,并利用隐私预算可组合性来保证整体隐私成本控制在预定值。我们将STEPS程序应用于模拟数据和2000-2012年当前人口调查青年选民数据。结果表明,与PrivBayes、改进的均匀直方图方法和平面拉普拉斯消毒方法相比,STEPS在某些分析中可以更好地保留种群水平信息和原始信息。
Differential Privacy (DP) formalizes privacy in mathematical terms and provides a robust concept for privacy protection. DIfferentially Private Data Synthesis (DIPS) techniques produce and release synthetic individual-level data in the DP framework. One key challenge to develop DIPS methods is the preservation of the statistical utility of synthetic data, especially in high-dimensional settings. We propose a new DIPS approach, STatistical Election to Partition Sequentially (STEPS) that partitions data by attributes according to their importance ranks according to either a practical or statistical importance measure. STEPS aims to achieve better original information preservation for the attributes with higher importance ranks and produce thus more useful synthetic data overall. We present an algorithm to implement the STEPS procedure and employ the privacy budget composability to ensure the overall privacy cost is controlled at the pre-specified value. We apply the STEPS procedure to both simulated data and the 2000-2012 Current Population Survey youth voter data. The results suggest STEPS can better preserve the population-level information and the original information for some analyses compared to PrivBayes, a modified Uniform histogram approach, and the flat Laplace sanitizer.