Autopopulus: A Novel Framework for Autoencoder Imputation on Large Clinical Datasets.

Autopopulus: A Novel Framework for Autoencoder Imputation on Large Clinical Datasets.
复制标题

DOI:
10.1109/embc46164.2021.9630135
复制
发表时间:
2021-11
期刊:
Annual International Conference of the IEEE Engineering in Medicine and Biology Society. IEEE Engineering in Medicine and Biology Society. Annual International Conference
影响因子:
--
通讯作者:
Sarrafzadeh M
Sarrafzadeh M
中科院分区:
其他
文献类型:
--
作者:
Zamanzadeh DJ;Petousis P;Davis TA;Nicholas SB;Norris KC;Tuttle KR;Bui AAT;Sarrafzadeh M

文献摘要

相似文献

电子健康记录(EHR)的采用使患者数据越来越容易获得,促进了各种临床决策支持系统和数据驱动模型的开发,以帮助医生。然而,缺失数据在EHR衍生的数据集中很常见,这可能会引入显著的不确定性,如果不是使预测模型的使用无效的话。基于机器学习(ML)的插补方法在各个领域都表现出了估计值和将不确定性降低到可以采用预测模型的程度的任务的前景。我们介绍了Autopopulus,一个新的框架,使各种自动编码器架构的设计和评估,有效地填补大型数据集。Autopopulus实现了现有的自动编码器方法以及一种新技术,该技术输出一系列估计值(而不是点估计值),并演示了一个工作流程,帮助用户对适当的插补方法做出明智的决定。为了进一步说明Autopopulus的实用性,我们不仅使用它来确定哪些插补方法可以最准确地对大型临床数据集进行插补,而且还确定了使下游预测模型能够实现预测慢性肾脏病(CKD)进展的最佳性能的插补方法。
The adoption of electronic health records (EHRs) has made patient data increasingly accessible, precipitating the development of various clinical decision support systems and data-driven models to help physicians. However, missing data are common in EHR-derived datasets, which can introduce significant uncertainty, if not invalidating the use of a predictive model. Machine learning (ML)-based imputation methods have shown promise in various domains for the task of estimating values and reducing uncertainty to the point that a predictive model can be employed. We introduce Autopopulus, a novel framework that enables the design and evaluation of various autoencoder architectures for efficient imputation on large datasets. Autopopulus implements existing autoencoder methods as well as a new technique that outputs a range of estimated values (rather than point estimates), and demonstrates a workflow that helps users make an informed decision on an appropriate imputation method. To further illustrate Autopopulus’ utility, we use it to identify not only which imputation methods can most accurately impute on a large clinical dataset, but to also identify the imputation methods that enable downstream predictive models to achieve the best performance for prediction of chronic kidney disease (CKD) progression.