Privacy-Preserving Generative Deep Neural Networks Support Clinical Data Sharing

Privacy-Preserving Generative Deep Neural Networks Support Clinical Data Sharing
复制标题

DOI:
10.1161/circoutcomes.118.005122
复制
发表时间:
2019-07-01
影响因子:
6.9
通讯作者:
Greene, Casey S.
Greene, Casey S.
中科院分区:
医学1区
文献类型:
--
作者:
Beaulieu-Jones, Brett K.;Wu, Zhiwei Steven;Greene, Casey S.

文献摘要

被引文献

相似文献

数据共享加速了科学进步,但在保护患者隐私的同时共享个人数据是一个障碍。方法和结果:使用成对的深度神经网络,我们生成了模拟的合成参与者,这些参与者与SPRINT试验(收缩压试验)的参与者非常相似。我们发现,这种配对网络可以用差分隐私进行训练,差分隐私是一种正式的隐私框架,它限制了对合成参与者数据的查询可以识别试验中真实的参与者的可能性。建立在合成群体上的机器学习预测器推广到原始数据集。这一发现表明,合成数据可以与其他人共享,使他们能够像拥有原始试验数据一样进行假设生成分析。结论:生成合成参与者的深度神经网络通过增强数据共享同时保护参与者隐私,促进了临床数据集的二次分析和可重复调查。
Background: Data sharing accelerates scientific progress but sharing individual-level data while preserving patient privacy presents a barrier. Methods and Results: Using pairs of deep neural networks, we generated simulated, synthetic participants that closely resemble participants of the SPRINT trial (Systolic Blood Pressure Trial). We showed that such paired networks can be trained with differential privacy, a formal privacy framework that limits the likelihood that queries of the synthetic participants' data could identify a real a participant in the trial. Machine learning predictors built on the synthetic population generalize to the original data set. This finding suggests that the synthetic data can be shared with others, enabling them to perform hypothesis-generating analyses as though they had the original trial data. Conclusions: Deep neural networks that generate synthetic participants facilitate secondary analyses and reproducible investigation of clinical data sets by enhancing data sharing while preserving participant privacy.