Clinical prediction in defined populations: a simulation study investigating when and how to aggregate existing models

Clinical prediction in defined populations: a simulation study investigating when and how to aggregate existing models
复制标题

特定人群的临床预测:一项模拟研究,调查何时以及如何聚合现有模型

DOI:
--
复制
发表时间:
2017
影响因子:
4
通讯作者:
M. Sperrin
M. Sperrin
中科院分区:
医学3区
文献类型:
--
作者:
G. Martin;M. Mamas;N. Peek;I. Buchan;M. Sperrin

文献摘要

参考文献

被引文献

相似文献

背景临床预测模型(CPM)越来越多地被用于支持医疗决策,但它们的推导并不一致,部分原因是数据有限。一种新的替代办法是汇总为类似环境和结果开发的现有国家方案管理方案。这项模拟研究的目的是通过建立一个新的模型来研究群体间的异质性和样本量对在一个定义的群体中聚集现有的CPM的影响。方法模拟被设计用来模拟这样一种场景,即在不同的、不同的群体中已经得到了两个结果的多个CPM,每个群体中都有潜在的不同的预测因子。然后,我们生成了一个新的“局部”群体,并通过聚合,使用堆叠回归、主成分分析或偏最小二乘,与从头开始使用反向选择和惩罚回归进行重新开发,比较了针对该群体开发的CPM的性能。结果虽然再开发方法导致模型对于少于500个观测值的本地数据集被错误校准,但模型聚合方法在所有模拟场景中都得到了很好的校准。当局部数据量小于1000个观测值且群体间异质性较小时,与建立新模型相比,聚合现有的CPM具有更好的区分性,并且预测风险的均方误差最小。相反,考虑到超过1000个观测值和显著的种群间异质性,重新开发的效果要好于聚集方法。在所有其他情景中,聚合和从头推导都会导致相似的预测性能。结论本研究展示了一种实用的方法来将CPM与定义的人群联系起来。当目标是在确定的人群中建立模型时,建模师应考虑现有的CPM,聚合方法是一种合适的建模策略,特别是在当地人口数据稀少的情况下。
BackgroundClinical prediction models (CPMs) are increasingly deployed to support healthcare decisions but they are derived inconsistently, in part due to limited data. An emerging alternative is to aggregate existing CPMs developed for similar settings and outcomes. This simulation study aimed to investigate the impact of between-population-heterogeneity and sample size on aggregating existing CPMs in a defined population, compared with developing a model de novo.MethodsSimulations were designed to mimic a scenario in which multiple CPMs for a binary outcome had been derived in distinct, heterogeneous populations, with potentially different predictors available in each. We then generated a new ‘local’ population and compared the performance of CPMs developed for this population by aggregation, using stacked regression, principal component analysis or partial least squares, with redevelopment from scratch using backwards selection and penalised regression.ResultsWhile redevelopment approaches resulted in models that were miscalibrated for local datasets of less than 500 observations, model aggregation methods were well calibrated across all simulation scenarios. When the size of local data was less than 1000 observations and between-population-heterogeneity was small, aggregating existing CPMs gave better discrimination and had the lowest mean square error in the predicted risks compared with deriving a new model. Conversely, given greater than 1000 observations and significant between-population-heterogeneity, then redevelopment outperformed the aggregation approaches. In all other scenarios, both aggregation and de novo derivation resulted in similar predictive performance.ConclusionThis study demonstrates a pragmatic approach to contextualising CPMs to defined populations. When aiming to develop models in defined populations, modellers should consider existing CPMs, with aggregation approaches being a suitable modelling strategy particularly with sparse data on the local population.
DOI: 10.1136/bmj.b605
发表时间: 2009-05-28
影响因子: 105.7
作者:
Altman, Douglas G.;Vergouwe, Yvonne;Moons, Karel G. M.
通讯作者: Moons, Karel G. M.