Private Data Synthesis from Decentralized Non-IID Data

Private Data Synthesis from Decentralized Non-IID Data
复制标题

DOI:
10.1109/ijcnn54540.2023.10191553
复制
发表时间:
2023-06
期刊:
2023 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Muhammad Usama Saleem;Liyue Fan
Muhammad Usama Saleem;Liyue Fan
中科院分区:
其他
文献类型:
--
作者:
Muhammad Usama Saleem;Liyue Fan

文献摘要

相似文献

保护隐私的数据共享可以进行广泛的探索性和二次数据分析,同时保护数据集中个人的隐私。机器学习的最新进展,特别是生成对抗网络(GAN),在合成真实数据集方面显示出了巨大的前景。在这项工作中,我们研究了在实际环境中私下训练 GAN 模型的可行性,其中输入数据分布在多方之间,并且本地数据可能高度倾斜,即非独立同分布。我们研究了应用于每个本地方的集中式私有 GAN 解决方案,并提出了一种提供强大隐私且适用于非 IID 数据的联合解决方案。我们对各种非独立同分布设置和来自不同领域的数据进行了广泛的实证分析。我们深入讨论了合成数据的效用、成员推理攻击方面的隐私风险,以及私有解决方案的隐私与效用权衡。
Privacy-preserving data sharing enables a wide range of exploratory and secondary data analyses while protecting the privacy of individuals in the datasets. Recent advancements in machine learning, specifically generative adversarial networks (GANs), have shown great promise for synthesizing realistic datasets. In this work, we investigate the feasibility of training GAN models privately in practical settings, where the input data is distributed across multiple parties, and local data may be highly skewed, i.e., non-IID. We examine centralized private GAN solutions applied at each local party and propose a federated solution that provides strong privacy and is suitable for non-IID data. We conduct extensive empirical analysis with a wide range of non-IID settings and data from different domains. We provide in-depth discussions about the utility of the synthetic data, the privacy risks in terms of membership inference attacks, as well as the privacy-utility trade-off for private solutions.