A comparison of approaches to improve worst-case predictive model performance over patient subpopulations.

A comparison of approaches to improve worst-case predictive model performance over patient subpopulations.
复制标题

DOI:
10.1038/s41598-022-07167-7
复制
发表时间:
2022-02-28
期刊:
影响因子:
4.6
通讯作者:
Shah NH
Shah NH
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Pfohl SR;Zhang H;Xu Y;Foryciarz A;Ghassemi M;Shah NH

文献摘要

参考文献

相似文献

在患者群体中平均准确的临床结果预测模型对某些亚群可能表现得很差,潜在地引入或加强了护理获取和质量方面的不平等。旨在最大化子群体中最坏情况模型性能的模型训练方法,如分布稳健优化(DRO),试图在不引入额外危害的情况下解决这一问题。我们对DRO和几种不同的标准学习程序进行了大规模的实证研究,以确定与从电子健康记录数据学习预测模型的标准方法相比,模型开发和选择的方法可以持续改善子群体的分类和最差情况下的性能。在我们的评估过程中,我们引入了对DRO方法的扩展,允许指定用于评估最坏情况性能的度量。我们对预测住院死亡率、住院时间延长和住院30天再入院的模型进行分析,并使用重症监护数据预测住院死亡率。我们发现,除了相对较少的例外,对于所检查的每个患者亚群,没有一种方法比使用整个训练数据集的标准学习程序执行得更好。这些结果暗示,当需要改进患者亚群的模型性能时,可能需要通过增加有效样本大小或降低预测问题中的噪声水平的数据收集技术来实现。
Predictive models for clinical outcomes that are accurate on average in a patient population may underperform drastically for some subpopulations, potentially introducing or reinforcing inequities in care access and quality. Model training approaches that aim to maximize worst-case model performance across subpopulations, such as distributionally robust optimization (DRO), attempt to address this problem without introducing additional harms. We conduct a large-scale empirical study of DRO and several variations of standard learning procedures to identify approaches for model development and selection that consistently improve disaggregated and worst-case performance over subpopulations compared to standard approaches for learning predictive models from electronic health records data. In the course of our evaluation, we introduce an extension to DRO approaches that allows for specification of the metric used to assess worst-case performance. We conduct the analysis for models that predict in-hospital mortality, prolonged length of stay, and 30-day readmission for inpatient admissions, and predict in-hospital mortality using intensive care data. We find that, with relatively few exceptions, no approach performs better, for each patient subpopulation examined, than standard learning procedures using the entire training dataset. These results imply that when it is of interest to improve model performance for patient subpopulations beyond what can be achieved with standard practices, it may be necessary to do so via data collection techniques that increase the effective sample size or reduce the level of noise in the prediction problem.
DOI: 10.1287/mnsc.1120.1641
发表时间: 2013-02-01
期刊: MANAGEMENT SCIENCE
影响因子: 5.4
作者:
Ben-Tal, Aharon;den Hertog, Dick;Rennen, Gijs
通讯作者: Rennen, Gijs
DOI: 10.1001/jamapsychiatry.2021.0493
发表时间: 2021-04-28
期刊: JAMA PSYCHIATRY
影响因子: 25.8
作者:
Coley, R. Yates;Johnson, Eric;Shortreed, Susan M.
通讯作者: Shortreed, Susan M.
DOI: 10.1093/jamia/ocaa283
发表时间: 2021-03-01
影响因子: 6.4
作者:
Barda, Noam;Yona, Gal;Dagan, Noa
通讯作者: Dagan, Noa
DOI: 10.1161/01.cir.101.23.e215
发表时间: 2000-06-13
期刊: CIRCULATION
影响因子: 37.8
作者:
Goldberger, AL;Amaral, LAN;Stanley, HE
通讯作者: Stanley, HE
DOI: 10.1002/sim.8281
发表时间: 2019-09-20
影响因子: 2
作者:
Austin, Peter C.;Steyerberg, Ewout W.
通讯作者: Steyerberg, Ewout W.