Machine Learning to Predict Mortality and Critical Events in a Cohort of Patients With COVID-19 in New York City: Model Development and Validation.

Machine Learning to Predict Mortality and Critical Events in a Cohort of Patients With COVID-19 in New York City: Model Development and Validation.
复制标题

DOI:
10.2196/24018
复制
发表时间:
2020-11-06
影响因子:
7.4
通讯作者:
Glicksberg BS
Glicksberg BS
中科院分区:
医学2区
文献类型:
--
作者:
Vaid A;Somani S;Russak AJ;De Freitas JK;Chaudhry FF;Paranjpe I;Johnson KW;Lee SJ;Miotto R;Richter F;Zhao S;Beckmann ND;Naik N;Kia A;Timsina P;Lala A;Paranjpe M;Golden E;Danieletto M;Singh M;Meyer D;O'Reilly PF;Huckins L;Kovatch P;Finkelstein J;Freeman RM;Argulian E;Kasarskis A;Percha B;Aberg JA;Bagiella E;Horowitz CR;Murphy B;Nestler EJ;Schadt EE;Cho JH;Cordon-Cardo C;Fuster V;Charney DS;Reich DL;Bottinger EP;Levin MA;Narula J;Fayad ZA;Just AC;Charney AW;Nadkarni GN;Glicksberg BS

文献摘要

参考文献

被引文献

相似文献

新冠肺炎已经感染了全球数以百万计的人,并导致数十万人死亡。新冠肺炎大流行需要深思熟虑的资源分配和早期识别高危患者。然而,缺乏有效的方法来满足这些需求。这项研究的目的是分析新冠肺炎检测呈阳性并进入纽约市西奈山医疗系统医院的患者的电子健康记录(EHR);开发机器学习模型,用于根据患者入院时的特征,在具有临床意义的时间范围内预测患者的医院病程;并评估这些模型在多家医院和时间点的性能。我们使用极端梯度增强(XGBoost)和基线比较模型来预测入院后3、5、7和10天的住院死亡率和危重事件。我们的研究人群包括来自纽约市五家医院的协调电子病历数据,其中包括2020年3月15日至5月22日收治的4,098名新冠肺炎阳性患者。这些模型首先在5月1日之前或当天在一家医院(n=1514)的患者身上进行培训,在5月1日之前或当天在其他四家医院(n=2201)的患者身上进行外部验证,并在5月1日之后的所有患者(n=383)上进行前瞻性验证。最后,我们建立了模型可解释性,以识别和排序驱动模型预测的变量。经交叉验证,XGBoost分类器优于基线模型,受试者工作特征曲线(AUC-ROC)下的面积为3天的死亡率为0.89,5天和7天的死亡率为0.85,10天的死亡率为0.84。XGBoost在关键事件预测方面也表现良好,AUC-ROC在3天时为0.80,5天时为0.79,7天时为0.80,10天时为0.81。在外部验证中,XGBoost在3天、5天、7天和10天预测死亡率的AUC-ROC分别为0.88、0.86、0.86和0.84。同样,未归结的XGBoost模型在3天时AUC-ROC为0.78,5天时为0.79,7天时为0.80,10天时为0.81。在预期验证集上的性能趋势相似。在第7天,入院时的急性肾损伤、LDH升高、呼吸急促和高血糖是危重事件预测的最强驱动因素,而高龄、阴离子间隙和C反应蛋白是死亡率预测的最强驱动因素。我们对新冠肺炎患者在不同时间段的死亡率和危重事件的机器学习模型进行了前瞻性的外部培训和验证。这些模型确定了高危患者,并揭示了预测结果的潜在关系。
COVID-19 has infected millions of people worldwide and is responsible for several hundred thousand fatalities. The COVID-19 pandemic has necessitated thoughtful resource allocation and early identification of high-risk patients. However, effective methods to meet these needs are lacking. The aims of this study were to analyze the electronic health records (EHRs) of patients who tested positive for COVID-19 and were admitted to hospitals in the Mount Sinai Health System in New York City; to develop machine learning models for making predictions about the hospital course of the patients over clinically meaningful time horizons based on patient characteristics at admission; and to assess the performance of these models at multiple hospitals and time points. We used Extreme Gradient Boosting (XGBoost) and baseline comparator models to predict in-hospital mortality and critical events at time windows of 3, 5, 7, and 10 days from admission. Our study population included harmonized EHR data from five hospitals in New York City for 4098 COVID-19–positive patients admitted from March 15 to May 22, 2020. The models were first trained on patients from a single hospital (n=1514) before or on May 1, externally validated on patients from four other hospitals (n=2201) before or on May 1, and prospectively validated on all patients after May 1 (n=383). Finally, we established model interpretability to identify and rank variables that drive model predictions. Upon cross-validation, the XGBoost classifier outperformed baseline models, with an area under the receiver operating characteristic curve (AUC-ROC) for mortality of 0.89 at 3 days, 0.85 at 5 and 7 days, and 0.84 at 10 days. XGBoost also performed well for critical event prediction, with an AUC-ROC of 0.80 at 3 days, 0.79 at 5 days, 0.80 at 7 days, and 0.81 at 10 days. In external validation, XGBoost achieved an AUC-ROC of 0.88 at 3 days, 0.86 at 5 days, 0.86 at 7 days, and 0.84 at 10 days for mortality prediction. Similarly, the unimputed XGBoost model achieved an AUC-ROC of 0.78 at 3 days, 0.79 at 5 days, 0.80 at 7 days, and 0.81 at 10 days. Trends in performance on prospective validation sets were similar. At 7 days, acute kidney injury on admission, elevated LDH, tachypnea, and hyperglycemia were the strongest drivers of critical event prediction, while higher age, anion gap, and C-reactive protein were the strongest drivers of mortality prediction. We externally and prospectively trained and validated machine learning models for mortality and critical events for patients with COVID-19 at different time horizons. These models identified at-risk patients and uncovered underlying relationships that predicted outcomes.
DOI: 10.1186/s12941-020-00362-2
发表时间: 2020-05-15
影响因子: 5.7
作者:
Chen, Wei;Zheng, Kenneth I.;Qiao, Zengpei
通讯作者: Qiao, Zengpei
DOI: 10.1016/j.jcv.2020.104370
发表时间: 2020-06-01
影响因子: 8.8
作者:
Liu, Fang;Li, Lin;Zhou, Xiang
通讯作者: Zhou, Xiang
DOI: 10.3390/jcm9061718
发表时间: 2020-06-01
影响因子: 3.9
作者:
Lim, Jeong-Hoon;Park, Sun-Hee;Kim, Yong-Lim
通讯作者: Kim, Yong-Lim
使用深度学习对重症 COVID-19 患者进行早期分诊
DOI: 10.1038/s41467-020-17280-8
发表时间: 2020-07-15
影响因子: 16.6
作者:
Liang, Wenhua;Yao, Jianhua;He, Jianxing
通讯作者: He, Jianxing
DOI: 10.2807/1560-7917.es.2020.25.10.2000180
发表时间: 2020-03-12
期刊: EUROSURVEILLANCE
影响因子: 19
作者:
Mizumoto, Kenji;Kagaya, Katsushi;Chowell, Gerardo
通讯作者: Chowell, Gerardo