Anonymization Through Data Synthesis Using Generative Adversarial Networks (ADS-GAN)

Anonymization Through Data Synthesis Using Generative Adversarial Networks (ADS-GAN)
复制标题

使用生成对抗网络(ADS-GAN)通过数据合成进行分析

DOI:
10.1109/jbhi.2020.2980262
复制
发表时间:
2020-08-01
影响因子:
7.7
通讯作者:
van der Schaar, Mihaela
van der Schaar, Mihaela
中科院分区:
工程技术1区
文献类型:
--
作者:
Yoon, Jinsung;Drumright, Lydia N.;van der Schaar, Mihaela

文献摘要

被引文献

相似文献

医疗和机器学习社区正在依靠人工智能 (AI) 的承诺,通过实现更准确的决策和个性化治疗来改变医学。然而,进展缓慢。未经同意的患者数据和隐私的法律和道德问题是数据共享的限制因素之一,导致机器学习社区访问常规收集的电子健康记录 (EHR) 时遇到重大障碍。我们提出了一种生成合成数据的新颖框架,该框架非常接近原始 EHR 数据集中变量的联合分布,提供易于访问、合法和道德上适当的解决方案来支持更开放的数据共享,从而促进人工智能解决方案的开发。为了解决充分匿名化定义不够清晰的问题,我们为“可识别性”创建了一个可量化的数学定义。我们使用条件生成对抗网络(GAN)框架来生成合成数据,同时最大限度地减少患者可识别性,该可识别性是根据给定任何个体患者的所有数据组合的重新识别概率来定义的。我们将适合我们综合生成的数据的模型与适合四个独立数据集的真实数据的模型进行比较,以评估模型性能的相似性,同时评估从合成数据中识别原始观察结果的程度。我们的模型 ADS-GAN 始终优于最先进的方法,并证明了联合分布的可靠性。我们建议这种方法可用于开发可以公开使用的数据集,同时大大降低泄露患者机密的风险。
The medical and machine learning communities are relying on the promise of artificial intelligence (AI) to transform medicine through enabling more accurate decisions and personalized treatment. However, progress is slow. Legal and ethical issues around unconsented patient data and privacy is one of the limiting factors in data sharing, resulting in a significant barrier in accessing routinely collected electronic health records (EHR) by the machine learning community. We propose a novel framework for generating synthetic data that closely approximates the joint distribution of variables in an original EHR dataset, providing a readily accessible, legally and ethically appropriate solution to support more open data sharing, enabling the development of AI solutions. In order to address issues around lack of clarity in defining sufficient anonymization, we created a quantifiable, mathematical definition for "identifiability". We used a conditional generative adversarial networks (GAN) framework to generate synthetic data while minimize patient identifiability that is defined based on the probability of re-identification given the combination of all data on any individual patient. We compared models fitted to our synthetically generated data to those fitted to the real data across four independent datasets to evaluate similarity in model performance, while assessing the extent to which original observations can be identified from the synthetic data. Our model, ADS-GAN, consistently outperformed state-of-the-art methods, and demonstrated reliability in the joint distributions. We propose that this method could be used to develop datasets that can be made publicly available while considerably lowering the risk of breaching patient confidentiality.