Democratizing EHR analyses with FIDDLE: a flexible data-driven preprocessing pipeline for structured clinical data.

Democratizing EHR analyses with FIDDLE: a flexible data-driven preprocessing pipeline for structured clinical data.
复制标题

使用FIDDLE民主化EHR分析:结构化临床数据的灵活数据驱动预处理管道。

DOI:
10.1093/jamia/ocaa139
复制
发表时间:
2020-12-09
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Wiens J
Wiens J
中科院分区:
其他
文献类型:
--
作者:
Tang S;Davarmanesh P;Song Y;Koutra D;Sjoding MW;Wiens J

文献摘要

参考文献

被引文献

相似文献

在将机器学习(ML)应用于电子健康记录(EHR)数据时,必须在应用任何ML之前做出许多决策;这种预处理需要大量工作,并且可能是劳动密集型的。随着ML在医疗保健中的作用越来越大,越来越需要针对EHR数据的系统化和可重复的预处理技术。因此,我们开发了FIDDLE(Flexible Data-Driven Pipeline),这是一个开源框架,可以简化从EHR中提取的数据的预处理。FIDDLE主要是数据驱动的,它系统地将结构化的EHR数据转换为特征向量,限制了用户必须做出的决策数量,同时结合了文献中的良好实践。为了证明它的实用性和灵活性,我们进行了一个概念验证实验,在实验中,我们将FIDDLE应用于从重症监护病房收集的2个公开可用的EHR数据集:MIMIC-III和eICU协作研究数据库。我们训练了不同的ML模型来预测3个临床上重要的结果:住院死亡率,急性呼吸衰竭和休克。我们使用受试者工作特征曲线下面积(AUROC)评估模型,并将其与几个基线进行比较。在所有任务中,FIDDLE分别从MIMIC-III和eICU中提取了2,528到7,403个特征。在所有任务中,基于FIDDLE的模型都取得了良好的区分性能,AUROC为0.757-0.886,与MIMIC-Extract的性能相当,MIMIC-Extract是专门为MIMIC-III设计的预处理管道。此外,我们的研究结果表明,FIDDLE在不同的预测时间,ML算法和数据集上是可推广的,同时对用户定义参数的不同设置具有相对鲁棒性。FIDDLE是一个开源预处理管道,有助于将ML应用于结构化EHR数据。通过加速和标准化劳动密集型预处理,FIDDLE可以帮助促进为EHR数据构建临床有用的ML工具的进展。
In applying machine learning (ML) to electronic health record (EHR) data, many decisions must be made before any ML is applied; such preprocessing requires substantial effort and can be labor-intensive. As the role of ML in health care grows, there is an increasing need for systematic and reproducible preprocessing techniques for EHR data. Thus, we developed FIDDLE (Flexible Data-Driven Pipeline), an open-source framework that streamlines the preprocessing of data extracted from the EHR. Largely data-driven, FIDDLE systematically transforms structured EHR data into feature vectors, limiting the number of decisions a user must make while incorporating good practices from the literature. To demonstrate its utility and flexibility, we conducted a proof-of-concept experiment in which we applied FIDDLE to 2 publicly available EHR data sets collected from intensive care units: MIMIC-III and the eICU Collaborative Research Database. We trained different ML models to predict 3 clinically important outcomes: in-hospital mortality, acute respiratory failure, and shock. We evaluated models using the area under the receiver operating characteristics curve (AUROC), and compared it to several baselines. Across tasks, FIDDLE extracted 2,528 to 7,403 features from MIMIC-III and eICU, respectively. On all tasks, FIDDLE-based models achieved good discriminative performance, with AUROCs of 0.757–0.886, comparable to the performance of MIMIC-Extract, a preprocessing pipeline designed specifically for MIMIC-III. Furthermore, our results showed that FIDDLE is generalizable across different prediction times, ML algorithms, and data sets, while being relatively robust to different settings of user-defined arguments. FIDDLE, an open-source preprocessing pipeline, facilitates applying ML to structured EHR data. By accelerating and standardizing labor-intensive preprocessing, FIDDLE can help stimulate progress in building clinically useful ML tools for EHR data.
DOI: 10.2196/medinform.5909
发表时间: 2016-09-30
影响因子: 3.2
作者:
Desautels T;Calvert J;Hoffman J;Jay M;Kerem Y;Shieh L;Shimabukuro D;Chettipally U;Feldman MD;Barton C;Wales DJ;Das R
通讯作者: Das R
DOI: 10.1126/scitranslmed.aab3719
发表时间: 2015-08-05
影响因子: 17.1
作者:
Henry, Katharine E.;Hager, David N.;Saria, Suchi
通讯作者: Saria, Suchi
DOI: 10.1093/ofid/ofz186
发表时间: 2019-05-01
影响因子: 4.2
作者:
Li, Benjamin Y.;Oh, Jeeheh;Wiens, Jenna
通讯作者: Wiens, Jenna
DOI: 10.1038/sdata.2016.35
发表时间: 2016-05-24
期刊: Scientific data
影响因子: 9.8
作者:
Johnson AE;Pollard TJ;Shen L;Lehman LW;Feng M;Ghassemi M;Moody B;Szolovits P;Celi LA;Mark RG
通讯作者: Mark RG
DOI: 10.1016/j.resuscitation.2016.02.005
发表时间: 2016-05-01
期刊: RESUSCITATION
影响因子: 6.5
作者:
Churpek, Matthew M.;Adhikari, Richa;Edelson, Dana P.
通讯作者: Edelson, Dana P.