Using Automated Machine Learning to Predict the Mortality of Patients With COVID-19: Prediction Model Development Study.

Using Automated Machine Learning to Predict the Mortality of Patients With COVID-19: Prediction Model Development Study.
复制标题

DOI:
10.2196/23458
复制
发表时间:
2021-02-26
影响因子:
7.4
通讯作者:
Reyes Gil M
Reyes Gil M
中科院分区:
医学2区
文献类型:
--
作者:
Ikemura K;Bellin E;Yagi Y;Billett H;Saada M;Simone K;Stahl L;Szymanski J;Goldstein DY;Reyes Gil M

文献摘要

参考文献

被引文献

相似文献

在大流行期间,临床医生对患者进行分层并决定谁获得有限的医疗资源非常重要。已经提出了机器学习模型来准确预测COVID-19疾病的严重程度。以前的研究通常只测试了一种机器学习算法,并将性能评估限制在曲线下面积分析。为了获得最佳结果,测试不同的机器学习算法以找到最佳预测模型可能很重要。在这项研究中,我们的目标是使用自动机器学习(autoML)来训练各种机器学习算法。我们选择了最能预测患者在SARS-CoV-2感染后存活几率的模型。此外,我们确定了哪些变量(即生命体征、生物标志物、合并症等)对生成准确模型最有影响。数据是从2020年3月1日至7月3日期间在我们机构检测出COVID-19呈阳性的所有患者中回顾性收集的。我们在指标时间(即实时聚合酶链反应阳性)之前或之后的36小时内收集了每位患者的48个变量。患者随访30天或直至死亡。患者的数据被用于通过autoML使用各种算法构建20个机器学习模型。机器学习模型的性能通过分析精确率-召回率曲线下面积(AUPCR)来衡量。随后,我们通过Shapley加法解释和部分依赖图建立了模型的可解释性,以识别和排名驱动模型预测的变量。之后,我们进行降维,提取10个最有影响力的变量。AutoML模型仅使用这10个变量进行重新训练,并针对使用48个变量的模型对输出模型进行评估。来自4313名患者的数据用于开发模型。使用autoML和48个变量生成的最佳模型是堆叠集合模型(AUPRC=0.807)。两个最好的独立模型是梯度增强机和极端梯度增强模型,其AUPRC分别为0.803和0.793。深度学习模型(AUPRC=0.73)明显劣于其他模型。对生成高性能模型最有影响力的10个变量是收缩压和舒张压、年龄、脉搏血氧饱和度水平、血尿素氮水平、乳酸脱氢酶水平、D-二聚体水平、肌钙蛋白水平、呼吸频率和Charlson合并症评分。在使用这10个变量重新训练autoML模型后,堆叠集成模型仍然具有最佳性能(AUPRC=0.791)。我们使用autoML开发了预测COVID-19患者生存率的高性能模型。此外,我们确定了与死亡率相关的重要变量。这证明了autoML是一种高效,有效和信息丰富的方法,用于生成基于机器学习的临床决策支持工具。
During a pandemic, it is important for clinicians to stratify patients and decide who receives limited medical resources. Machine learning models have been proposed to accurately predict COVID-19 disease severity. Previous studies have typically tested only one machine learning algorithm and limited performance evaluation to area under the curve analysis. To obtain the best results possible, it may be important to test different machine learning algorithms to find the best prediction model. In this study, we aimed to use automated machine learning (autoML) to train various machine learning algorithms. We selected the model that best predicted patients’ chances of surviving a SARS-CoV-2 infection. In addition, we identified which variables (ie, vital signs, biomarkers, comorbidities, etc) were the most influential in generating an accurate model. Data were retrospectively collected from all patients who tested positive for COVID-19 at our institution between March 1 and July 3, 2020. We collected 48 variables from each patient within 36 hours before or after the index time (ie, real-time polymerase chain reaction positivity). Patients were followed for 30 days or until death. Patients’ data were used to build 20 machine learning models with various algorithms via autoML. The performance of machine learning models was measured by analyzing the area under the precision-recall curve (AUPCR). Subsequently, we established model interpretability via Shapley additive explanation and partial dependence plots to identify and rank variables that drove model predictions. Afterward, we conducted dimensionality reduction to extract the 10 most influential variables. AutoML models were retrained by only using these 10 variables, and the output models were evaluated against the model that used 48 variables. Data from 4313 patients were used to develop the models. The best model that was generated by using autoML and 48 variables was the stacked ensemble model (AUPRC=0.807). The two best independent models were the gradient boost machine and extreme gradient boost models, which had an AUPRC of 0.803 and 0.793, respectively. The deep learning model (AUPRC=0.73) was substantially inferior to the other models. The 10 most influential variables for generating high-performing models were systolic and diastolic blood pressure, age, pulse oximetry level, blood urea nitrogen level, lactate dehydrogenase level, D-dimer level, troponin level, respiratory rate, and Charlson comorbidity score. After the autoML models were retrained with these 10 variables, the stacked ensemble model still had the best performance (AUPRC=0.791). We used autoML to develop high-performing models that predicted the survival of patients with COVID-19. In addition, we identified important variables that correlated with mortality. This is proof of concept that autoML is an efficient, effective, and informative method for generating machine learning–based clinical decision support tools.
DOI: 10.1001/jamanetworkopen.2020.23934
发表时间: 2020-10-01
期刊: JAMA network open
影响因子: 13.8
作者:
Castro VM;McCoy TH;Perlis RH
通讯作者: Perlis RH
COVID-19 肺炎患者的肾脏受累和早期预后
DOI: 10.1681/asn.2020030276
发表时间: 2020-06-01
影响因子: 13.6
作者:
Pei, Guangchang;Zhang, Zhiguo;Xu, Gang
通讯作者: Xu, Gang
DOI: 10.1214/aos/1013203451
发表时间: 2001-10-01
影响因子: 4.5
作者:
Friedman, JH
通讯作者: Friedman, JH
DOI: 10.3390/jcm9061668
发表时间: 2020-06-01
影响因子: 3.9
作者:
Cheng, Fu-Yuan;Joshi, Himanshu;Kia, Arash
通讯作者: Kia, Arash
DOI: 10.2196/24018
发表时间: 2020-11-06
影响因子: 7.4
作者:
Vaid A;Somani S;Russak AJ;De Freitas JK;Chaudhry FF;Paranjpe I;Johnson KW;Lee SJ;Miotto R;Richter F;Zhao S;Beckmann ND;Naik N;Kia A;Timsina P;Lala A;Paranjpe M;Golden E;Danieletto M;Singh M;Meyer D;O'Reilly PF;Huckins L;Kovatch P;Finkelstein J;Freeman RM;Argulian E;Kasarskis A;Percha B;Aberg JA;Bagiella E;Horowitz CR;Murphy B;Nestler EJ;Schadt EE;Cho JH;Cordon-Cardo C;Fuster V;Charney DS;Reich DL;Bottinger EP;Levin MA;Narula J;Fayad ZA;Just AC;Charney AW;Nadkarni GN;Glicksberg BS
通讯作者: Glicksberg BS