Safe and reliable transport of prediction models to new healthcare settings without the need to collect new labeled data.

Safe and reliable transport of prediction models to new healthcare settings without the need to collect new labeled data.
复制标题

将预测模型安全可靠地传输到新的医疗保健环境,而无需收集新的标记数据。

DOI:
10.1101/2023.12.13.23299899
复制
发表时间:
2023
期刊:
medRxiv : the preprint server for health sciences
影响因子:
--
通讯作者:
Beam,Andrew
Beam,Andrew
中科院分区:
--
文献类型:
--
作者:
Tuwani,Rudraksh;Beam,Andrew

文献摘要

相似文献

从业者和临床医生如何知道在不同机构训练的预测模型是否可以安全地用于他们的患者人群?有大量证据表明,预测模型使用的协变量分布的微小变化可能导致它们在部署到新设置时失败。这种特定类型的数据集转移,称为协变量转移,是在新的医疗保健环境中实施现有预测模型的核心挑战。一种解决方案是在目标人群中收集额外的标签,然后微调预测模型,使其适应新的医疗环境的特征,这通常被称为本地化。然而,收集新的标签可能是昂贵和耗时的。为了解决这些问题,我们在不确定性量化方面重新定义了模型传输的核心问题,这使得人们可以知道在一个环境中训练的模型何时可以安全地用于新的医疗环境。使用保角预测的方法,我们展示了如何在存在协变量偏移的情况下在不同设置之间安全地传输模型,即使所有人都可以访问来自新设置的协变量(例如,没有新标签)。使用这种方法,模型返回一个量化其不确定性的预测集,并保证包含具有用户指定概率(例如90%)的正确标签,该属性也称为覆盖率。我们表明,加权共形推理过程的基础上的源和目标人群之间的密度比估计可以产生预测集与真实世界的数据覆盖的正确水平。这使用户能够知道模型的预测是否可以信任,而无需收集新的标记数据。
How can practitioners and clinicians know if a prediction model trained at a different institution can be safely used on their patient population? There is a large body of evidence showing that small changes in the distribution of the covariates used by prediction models may cause them to fail when deployed to new settings. This specific kind of dataset shift, known as covariate shift, is a central challenge to implementing existing prediction models in new healthcare environments. One solution is to collect additional labels in the target population and then fine tune the prediction model to adapt it to the characteristics of the new healthcare setting, which is often referred to as localization. However, collecting new labels can be expensive and time-consuming. To address these issues, we recast the core problem of model transportation in terms of uncertainty quantification, which allows one to know when a model trained in one setting may be safely used in a new healthcare environment of interest. Using methods from conformal prediction, we show how to transport models safely between different settings in the presence of covariate shift, even when all one has access to are covariates from the new setting of interest (e.g. no new labels). Using this approach, the model returns a prediction set that quantifies its uncertainty and is guaranteed to contain the correct label with a user-specified probability (e.g. 90%), a property that is also known as coverage. We show that a weighted conformal inference procedure based on density ratio estimation between the source and target populations can produce prediction sets with the correct level of coverage on real-world data. This allows users to know if a model’s predictions can be trusted on their population without the need to collect new labeled data.