Generating high-fidelity synthetic patient data for assessing machine learning healthcare software.

Generating high-fidelity synthetic patient data for assessing machine learning healthcare software.
复制标题

DOI:
10.1038/s41746-020-00353-9
复制
发表时间:
2020-11-09
影响因子:
15.2
通讯作者:
Myles P
Myles P
中科院分区:
医学1区
文献类型:
--
作者:
Tucker A;Wang Z;Rotalinti Y;Myles P

文献摘要

参考文献

被引文献

相似文献

在医疗保健系统中采用现代人工智能技术的需求越来越大。其中许多技术利用历史患者健康数据来构建强大的预测模型,可用于改善对疾病的诊断和理解。然而,为了使所有部门能够更好地利用这些数据,需要考虑许多有关患者隐私的问题。一种可以提供规避隐私问题的方法的方法是创建真实的合成数据集,其捕获与原始数据集一样多的复杂性(分布、非线性关系和噪声),但实际上不包括任何真实的患者数据。虽然以前的研究已经探索了生成合成数据集的模型,但在这里,我们将探索基于英国初级保健患者数据的RESISTANCE,概率图形建模,潜在变量识别和离群值分析的集成,以生成逼真的合成数据。特别是,我们专注于处理缺失,变量之间的复杂相互作用,以及机器学习分类器产生的敏感性分析统计数据,同时量化从合成数据点重新识别患者的风险。我们表明,通过我们将离群值分析与图形建模和恢复相结合的方法,我们可以实现在推断机器学习分类器时,在特征分布,特征依赖性和敏感性分析统计方面与原始地面真实数据没有显着差异的合成数据集。更重要的是,生成与真实的患者相同或非常相似的合成数据的风险被证明是低的。
There is a growing demand for the uptake of modern artificial intelligence technologies within healthcare systems. Many of these technologies exploit historical patient health data to build powerful predictive models that can be used to improve diagnosis and understanding of disease. However, there are many issues concerning patient privacy that need to be accounted for in order to enable this data to be better harnessed by all sectors. One approach that could offer a method of circumventing privacy issues is the creation of realistic synthetic data sets that capture as many of the complexities of the original data set (distributions, non-linear relationships, and noise) but that does not actually include any real patient data. While previous research has explored models for generating synthetic data sets, here we explore the integration of resampling, probabilistic graphical modelling, latent variable identification, and outlier analysis for producing realistic synthetic data based on UK primary care patient data. In particular, we focus on handling missingness, complex interactions between variables, and the resulting sensitivity analysis statistics from machine learning classifiers, while quantifying the risks of patient re-identification from synthetic datapoints. We show that, through our approach of integrating outlier analysis with graphical modelling and resampling, we can achieve synthetic data sets that are not significantly different from original ground truth data in terms of feature distributions, feature dependencies, and sensitivity analysis statistics when inferring machine learning classifiers. What is more, the risk of generating synthetic data that is identical or very similar to real patients is shown to be low.
DOI: 10.1007/s10194-010-0282-4
发表时间: 2011-04
影响因子: 7.4
作者:
Antonaci, Fabio;Nappi, Giuseppe;Galli, Federica;Manzoni, Gian Camillo;Calabresi, Paolo;Costa, Alfredo
通讯作者: Costa, Alfredo
DOI: 10.1016/j.hfc.2008.03.008
发表时间: 2008-10
影响因子: 3.4
作者:
Ahmed, Ali;Campbell, Ruth C
通讯作者: Campbell, Ruth C
DOI: 10.1177/0962280214558972
发表时间: 2017-04
影响因子: 2.3
作者:
Austin PC;Steyerberg EW
通讯作者: Steyerberg EW
DOI: 10.1136/bmj.39609.449676.25
发表时间: 2008-06-28
影响因子: 105.7
作者:
Hippisley-Cox, Julia;Coupland, Carol;Brindle, Peter
通讯作者: Brindle, Peter
DOI: 10.1111/j.2517-6161.1977.tb01600.x
发表时间: 1977-01-01
期刊: JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-METHODOLOGICAL
影响因子: --
作者:
DEMPSTER, AP;LAIRD, NM;RUBIN, DB
通讯作者: RUBIN, DB