Evaluating Identity Disclosure Risk in Fully Synthetic Health Data: Model Development and Validation.

Evaluating Identity Disclosure Risk in Fully Synthetic Health Data: Model Development and Validation.
复制标题

评估全合成健康数据中的身份披露风险:模型开发和验证。

DOI:
10.2196/23139
复制
发表时间:
2020-11-16
影响因子:
7.4
通讯作者:
Bass J
Bass J
中科院分区:
医学2区
文献类型:
--
作者:
El Emam K;Mosquera L;Bass J

文献摘要

参考文献

被引文献

相似文献

人们对数据综合越来越感兴趣,以便能够共享用于二级分析的数据;然而,对于完全合成的数据,需要一个全面的隐私风险模型:如果生成模型已经过拟合,那么就有可能从合成数据中识别个人并了解他们的新信息。本研究的目的是开发并应用一种方法来评估完全合成数据的身份披露风险。提出了一个完整的风险模型,该模型既评估了身份披露,也评估了攻击者在合成记录与真人匹配的情况下学习新东西的能力。我们称之为“有意义的身份披露风险”。该模型应用于华盛顿州医院出院数据库(2007年)和加拿大COVID-19病例数据库的样本。这两个数据集是使用通常用于合成健康和社会科学数据的顺序决策树过程合成的。这两个合成样本的有意义身份披露风险均低于常用的0.09风险阈值(分别为0.0198和0.0086),分别比原始数据集的风险值低4倍和5倍。我们提出了一个完整合成数据的身份披露风险模型。在2个数据集上的结果表明,这种综合方法可以显著降低有意义的身份披露风险。该风险模型可用于全合成数据的隐私性评估。
There has been growing interest in data synthesis for enabling the sharing of data for secondary analysis; however, there is a need for a comprehensive privacy risk model for fully synthetic data: If the generative models have been overfit, then it is possible to identify individuals from synthetic data and learn something new about them. The purpose of this study is to develop and apply a methodology for evaluating the identity disclosure risks of fully synthetic data. A full risk model is presented, which evaluates both identity disclosure and the ability of an adversary to learn something new if there is a match between a synthetic record and a real person. We term this “meaningful identity disclosure risk.” The model is applied on samples from the Washington State Hospital discharge database (2007) and the Canadian COVID-19 cases database. Both of these datasets were synthesized using a sequential decision tree process commonly used to synthesize health and social science data. The meaningful identity disclosure risk for both of these synthesized samples was below the commonly used 0.09 risk threshold (0.0198 and 0.0086, respectively), and 4 times and 5 times lower than the risk values for the original datasets, respectively. We have presented a comprehensive identity disclosure risk model for fully synthetic data. The results for this synthesis method on 2 datasets demonstrate that synthesis can reduce meaningful identity disclosure risks considerably. The risk model can be applied in the future to evaluate the privacy of fully synthetic data.
DOI: 10.1038/srep01376
发表时间: 2013
期刊: SCIENTIFIC REPORTS
影响因子: 4.6
作者:
de Montjoye, Yves-Alexandre;Hidalgo, Cesar A.;Verleysen, Michel;Blondel, Vincent D.
通讯作者: Blondel, Vincent D.
DOI: 10.1037/pspp0000208
发表时间: 2021-08-01
影响因子: 7.6
作者:
Arslan, Ruben C.;Schilling, Katharina M.;Penke, Lars
通讯作者: Penke, Lars
DOI: 10.1038/ng.142
发表时间: 2008-05-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Blewitt, Marnie E.;Gendrel, Anne-Valerie;Whitelaw, Emma
通讯作者: Whitelaw, Emma
DOI: 10.1136/jamia.2009.000026
发表时间: 2010-03-01
影响因子: 6.4
作者:
Benitez, Kathleen;Malin, Bradley
通讯作者: Malin, Bradley
DOI: 10.1080/19345747.2019.1631421
发表时间: 2019-07-22
影响因子: 1.8
作者:
Bonnery, Daniel;Feng, Yi;Zheng, Yating
通讯作者: Zheng, Yating