Evaluating the utility of synthetic COVID-19 case data.

Evaluating the utility of synthetic COVID-19 case data.
复制标题

DOI:
10.1093/jamiaopen/ooab012
复制
发表时间:
2021-01
期刊:
影响因子:
2.1
通讯作者:
Sood H
Sood H
中科院分区:
其他
文献类型:
--
作者:
El Emam K;Mosquera L;Jonker E;Sood H

文献摘要

参考文献

被引文献

相似文献

对患者隐私的担忧限制了对COVID-19数据集的访问。数据综合是以保护隐私的方式向研究界广泛提供此类数据的一种方法。通过比较真实数据和合成数据的分析结果来评估合成数据的效用。利用安大略省90514例与社区合并症、人口统计学和社会经济特征相关的COVID-19病例记录,建立了一个梯度增强分类树来预测死亡。评估了模型的准确性和关系,以及隐私风险。在合成数据集上开发了相同的模型,并与原始数据的模型进行了比较。真实数据模型的AUROC和AUPRC分别为0.945[95%可信区间(CI), 0.941 ~ 0.948]和0.34 (95% CI, 0.313 ~ 0.368)。综合数据模型的AUROC和AUPRC分别为0.94 (95% CI, 0.936-0.944)和0.313 (95% CI, 0.286-0.342),置信区间与真实数据重叠度分别为45.05%和52.02%。真实模型和合成模型中最重要的死亡预测因子按降序排列:年龄、自2020年1月1日以来的天数、暴露类型和性别。两个数据集之间的函数关系相似。属性披露风险为0.0585,成员披露风险较低。这个合成数据集可以用作真实数据集的代理。
Concerns about patient privacy have limited access to COVID-19 datasets. Data synthesis is one approach for making such data broadly available to the research community in a privacy protective manner. Evaluate the utility of synthetic data by comparing analysis results between real and synthetic data. A gradient boosted classification tree was built to predict death using Ontario’s 90 514 COVID-19 case records linked with community comorbidity, demographic, and socioeconomic characteristics. Model accuracy and relationships were evaluated, as well as privacy risks. The same model was developed on a synthesized dataset and compared to one from the original data. The AUROC and AUPRC for the real data model were 0.945 [95% confidence interval (CI), 0.941–0.948] and 0.34 (95% CI, 0.313–0.368), respectively. The synthetic data model had AUROC and AUPRC of 0.94 (95% CI, 0.936–0.944) and 0.313 (95% CI, 0.286–0.342) with confidence interval overlap of 45.05% and 52.02% when compared with the real data. The most important predictors of death for the real and synthetic models were in descending order: age, days since January 1, 2020, type of exposure, and gender. The functional relationships were similar between the two data sets. Attribute disclosure risks were 0.0585, and membership disclosure risk was low. This synthetic dataset could be used as a proxy for the real dataset.
DOI: 10.1093/jamia/ocaa249
发表时间: 2021-01-15
期刊: Journal of the American Medical Informatics Association : JAMIA
影响因子: --
作者:
Emam KE;Mosquera L;Zheng C
通讯作者: Zheng C
DOI: 10.1186/1471-2458-11-454
发表时间: 2011-06-09
期刊: BMC public health
影响因子: 4.5
作者:
El Emam K;Mercer J;Moreau K;Grava-Gubins I;Buckeridge D;Jonker E
通讯作者: Jonker E
DOI: 10.2196/23139
发表时间: 2020-11-16
影响因子: 7.4
作者:
El Emam K;Mosquera L;Bass J
通讯作者: Bass J
DOI: 10.1038/s41467-020-18297-9
发表时间: 2020-09-07
影响因子: 16.6
作者:
Barda, Noam;Riesel, Dan;Dagan, Noa
通讯作者: Dagan, Noa
DOI: 10.1037/pspp0000208
发表时间: 2021-08-01
影响因子: 7.6
作者:
Arslan, Ruben C.;Schilling, Katharina M.;Penke, Lars
通讯作者: Penke, Lars