Utility Analysis of Horizontally Merged Multi-Party Synthetic Data with Differential Privacy

Utility Analysis of Horizontally Merged Multi-Party Synthetic Data with Differential Privacy
复制标题

DOI:
10.1109/isncc49221.2020.9297254
复制
发表时间:
2020-10
期刊:
2020 International Symposium on Networks, Computers and Communications (ISNCC)
影响因子:
--
通讯作者:
Bingyue Su;Fang Liu
Bingyue Su;Fang Liu
中科院分区:
其他
文献类型:
--
作者:
Bingyue Su;Fang Liu

文献摘要

相似文献

通常需要大量的数据来训练机器学习算法。实现必要数据量的一种方法是共享和联合收割机来自多方的数据。另一方面,如何在数据共享过程中保护敏感的个人信息始终是一个挑战。我们专注于数据共享时,各方有重叠的属性,但不重叠的个人。实现隐私保护的一种方法是通过共享不同的私有合成数据。每一方都以自己偏好的隐私预算生成合成数据,然后发布并在各方之间横向合并。这种方法的总隐私成本以一方所使用的最大个人预算为上限。我们推导出的均方误差界的参数估计在常见的回归分析的基础上合并消毒的数据跨越各方。我们通过理论分析确定的条件下,共享和合并消毒数据的效用超过扰动引入满足差分隐私和超越基于个人数据。实验表明,以实际合理的小隐私成本获得的经过消毒的HOMM数据可以导致比单个方更小的预测和估计误差,这表明了数据共享的好处,同时保护隐私。
A large amount of data is often needed to train machine learning algorithms with confidence. One way to achieve the necessary data volume is to share and combine data from multiple parties. On the other hand, how to protect sensitive personal information during data sharing is always a challenge. We focus on data sharing when parties have overlapping attributes but non-overlapping individuals. One approach to achieve privacy protection is through sharing differentially private synthetic data. Each party generates synthetic data at its own preferred privacy budget, which are then released and horizontally merged across the parties. The total privacy cost for this approach is capped at the maximum individual budget employed by a party. We derive the mean squared error bounds for the parameter estimation in common regression analysis based on the merged sanitized data across parties. We identify through theoretical analysis the conditions under which the utility of sharing and merging sanitized data outweighs the perturbation introduced for satisfying differential privacy and surpasses that based on individual party data. The experiments suggest that sanitized HOMM data obtained at a practically reasonable small privacy cost can lead to smaller prediction and estimation errors than individual parties, demonstrating benefits of data sharing while protecting privacy.