Machine learning for patient risk stratification: standing on, or looking over, the shoulders of clinicians?

Machine learning for patient risk stratification: standing on, or looking over, the shoulders of clinicians?
复制标题

DOI:
10.1038/s41746-021-00426-3
复制
发表时间:
2021-03-30
影响因子:
15.2
通讯作者:
Kohane IS
Kohane IS
中科院分区:
医学1区
文献类型:
--
作者:
Beaulieu-Jones BK;Yuan W;Brat GA;Beam AL;Weber G;Ruffin M;Kohane IS

文献摘要

参考文献

被引文献

相似文献

只有当研究人员展示了能够提供新颖见解的模型,而不是学习临床医生将采取的一系列行动中最有可能采取的下一步行动时,机器学习才能帮助临床医生做出个性化的患者预测。我们仅使用临床医生发起的4290万次入院的行政数据训练深度学习模型,使用三个数据子集:仅人口统计数据,入院时可用的人口统计数据和信息,以及入院第一天记录的先前数据和收费。在入院第一天进行收费训练的模型的表现接近已公布的基于emr的住院结果基准:住院死亡率(0.89 AUC)、住院时间延长(0.82 AUC)和30天再入院率(0.71 AUC)。仅使用临床启动数据训练的模型与使用完整电子病历数据训练的模型(据称包括患者状态和生理信息)之间的相似表现应引起对这些模型部署的关注。此外,与仅对心肌梗死(MI)患者进行训练的模型相比,这些模型在仅对心肌梗死(MI)患者进行评估时表现出显著的性能下降,这突出了医生诊断对这些模型预后性能的重要性。这些结果为仅根据先前临床行为训练的预测准确性提供了基准,并表明具有类似性能的模型可能通过观察临床医生的肩膀来获得信号-使用临床行为作为先前存在的直觉和怀疑的表达来生成预测。对于指导临床医生进行个人决策的模型来说,超过这些基准的性能是必要的。
Machine learning can help clinicians to make individualized patient predictions only if researchers demonstrate models that contribute novel insights, rather than learning the most likely next step in a set of actions a clinician will take. We trained deep learning models using only clinician-initiated, administrative data for 42.9 million admissions using three subsets of data: demographic data only, demographic data and information available at admission, and the previous data plus charges recorded during the first day of admission. Models trained on charges during the first day of admission achieve performance close to published full EMR-based benchmarks for inpatient outcomes: inhospital mortality (0.89 AUC), prolonged length of stay (0.82 AUC), and 30-day readmission rate (0.71 AUC). Similar performance between models trained with only clinician-initiated data and those trained with full EMR data purporting to include information about patient state and physiology should raise concern in the deployment of these models. Furthermore, these models exhibited significant declines in performance when evaluated over only myocardial infarction (MI) patients relative to models trained over MI patients alone, highlighting the importance of physician diagnosis in the prognostic performance of these models. These results provide a benchmark for predictive accuracy trained only on prior clinical actions and indicate that models with similar performance may derive their signal by looking over clinician’s shoulders—using clinical behavior as the expression of preexisting intuition and suspicion to generate a prediction. For models to guide clinicians in individual decisions, performance exceeding these benchmarks is necessary.
DOI: 10.1093/jamia/ocw054
发表时间: 2017-01-01
影响因子: 6.4
作者:
van der Bij, Sjoukje;Khan, Nasra;Verheij, Robert A.
通讯作者: Verheij, Robert A.
DOI: 10.1038/s41746-018-0029-1
发表时间: 2018-05-08
影响因子: 15.2
作者:
Rajkomar, Alvin;Oren, Eyal;Dean, Jeffrey
通讯作者: Dean, Jeffrey
DOI: 10.1377/hlthaff.2014.0038
发表时间: 2014-07-01
期刊: HEALTH AFFAIRS
影响因子: 9.7
作者:
Wallace, Paul J.;Shah, Nilay D.;Crown, William H.
通讯作者: Crown, William H.
DOI: 10.1038/s41746-021-00426-3
发表时间: 2021-03-30
影响因子: 15.2
作者:
Beaulieu-Jones BK;Yuan W;Brat GA;Beam AL;Weber G;Ruffin M;Kohane IS
通讯作者: Kohane IS
DOI: 10.1136/bmj.k1479
发表时间: 2018-04-30
期刊: BMJ (Clinical research ed.)
影响因子: --
作者:
Agniel D;Kohane IS;Weber GM
通讯作者: Weber GM