Data Synthesis based on Generative Adversarial Networks

Data Synthesis based on Generative Adversarial Networks
复制标题

DOI:
10.14778/3231751.3231757
复制
发表时间:
2018-06-01
影响因子:
2.5
通讯作者:
Kim, Youngmin
Kim, Youngmin
中科院分区:
计算机科学2区
文献类型:
--
作者:
Park, Noseong;Mohammadi, Mahmoud;Kim, Youngmin

文献摘要

被引文献

相似文献

隐私是我们社会的一个重要问题,与合作伙伴共享数据或向公众发布数据是经常发生的事情。一些用于实现隐私的技术是删除标识符,更改准标识符和扰动值。不幸的是,这些方法受到两个限制。首先,已经表明,如果攻击者拥有一些背景知识或其他信息源,私人信息仍然可以泄露。其次,它们没有考虑到这些方法对所发布数据的效用的不利影响。在本文中,我们提出了一种方法,满足这两个要求。我们的方法称为table-GAN,使用生成对抗网络(GAN)来合成与原始表在统计上相似但不会导致信息泄漏的假表。我们表明,使用我们的合成表训练的机器学习模型表现出的性能与使用原始表训练的未知测试用例的模型相似。我们称之为属性模型兼容性。我们认为,没有模型兼容性的匿名化/扰动/合成方法是没有价值的。我们使用了来自四个不同领域的四个真实世界数据集进行实验,并与最先进的匿名化,扰动和生成技术进行了深入的比较。在我们的实验中,只有我们的方法始终显示隐私级别和模型兼容性之间的平衡。
Privacy is an important concern for our society where sharing data with partners or releasing data to the public is a frequent occurrence. Some of the techniques that are being used to achieve privacy are to remove identifiers, alter quasi-identifiers, and perturb values. Unfortunately, these approaches suffer from two limitations. First, it has been shown that private information can still be leaked if attackers possess some background knowledge or other information sources. Second, they do not take into account the adverse impact these methods will have on the utility of the released data. In this paper, we propose a method that meets both requirements. Our method, called table-GAN, uses generative adversarial networks (GANs) to synthesize fake tables that are statistically similar to the original table yet do not incur information leakage. We show that the machine learning models trained using our synthetic tables exhibit performance that is similar to that of models trained using the original table for unknown testing cases. We call this property model compatibility. We believe that anonymization/perturbation/synthesis methods without model compatibility are of little value. We used four real-world datasets from four different domains for our experiments and conducted in-depth comparisons with state-of-the-art anonymization, perturbation, and generation techniques. Throughout our experiments, only our method consistently shows balance between privacy level and model compatibility.