Data Synthesis via Differentially Private Markov Random Fields

Data Synthesis via Differentially Private Markov Random Fields
复制标题

DOI:
10.14778/3476249.3476272
复制
发表时间:
2021-07-01
影响因子:
2.5
通讯作者:
Xiao, Xiaokui
Xiao, Xiaokui
中科院分区:
计算机科学2区
文献类型:
--
作者:
Cai, Kuntai;Lei, Xiaoyu;Xiao, Xiaokui

文献摘要

被引文献

相似文献

本文研究了具有差分隐私(DP)的高维数据集的合成问题。现有技术的解决方案通过首先生成输入数据的噪声低维边缘的集合M来解决这个问题。然后用它们来近似数据分布在...用于合成数据生成。然而,它对M施加了几个限制,大大限制了边缘的选择。这使得它很难捕捉所有重要的属性之间的相关性,这反过来又降低了合成data.To解决上述缺陷的质量,我们提出PrivMRF,一种方法,(i)也利用了一套M的低维边缘合成高维数据与DP,但(ii)提供了高度的灵活性,在边缘的选择。PrivMRF的核心思想是选择一个合适的M来构造一个马尔可夫随机场(MRF)来模拟输入数据中属性之间的相关性,然后使用MRF进行数据综合。在四个基准数据集上的实验结果表明,PrivMRF在对生成的合成数据进行计数查询和分类任务的准确性方面始终优于现有技术。
This paper studies the synthesis of high-dimensional datasets with differential privacy (DP). The state-of-the-art solution addresses this problem by first generating a setM of noisy low-dimensional marginals of the input data.., and then use them to approximate the data distribution in.. for synthetic data generation. However, it imposes several constraints on M that considerably limits the choices of marginals. This makes it difficult to capture all important correlations among attributes, which in turn degrades the quality of the resulting synthetic data.To address the above deficiency, we propose PrivMRF, a method that (i) also utilizes a setM of low-dimensional marginals for synthesizing high-dimensional data with DP, but (ii) provides a high degree of flexibility in the choices of marginals. The key idea of PrivMRF is to select an appropriateM to construct a Markov random field (MRF) that models the correlations among the attributes in the input data, and then use the MRF for data synthesis. Experimental results on four benchmark datasets show that PrivMRF consistently outperforms the state of the art in terms of the accuracy of counting queries and classification tasks conducted on the synthetic data generated.