Minimum sample size for external validation of a clinical prediction model with a continuous outcome

Minimum sample size for external validation of a clinical prediction model with a continuous outcome
复制标题

DOI:
10.1002/sim.8766
复制
发表时间:
2020-11-04
影响因子:
2
通讯作者:
Riley, Richard D.
Riley, Richard D.
中科院分区:
医学3区
文献类型:
--
作者:
Archer, Lucinda;Snell, Kym I. E.;Riley, Richard D.

文献摘要

被引文献

相似文献

临床预测模型提供个性化的结果预测,为患者咨询和临床决策提供信息。外部验证是在独立于用于模型开发的数据中检查预测模型的性能的过程。目前的外部验证研究通常存在样本量小,因此对模型预测性能的估计不准确的问题。为了解决这一问题,我们建议如何确定临床预测模型的外部验证所需的最小样本量,并获得连续的结果。提出了四个标准,即目标精确估计数(1)R-2(解释的方差比例)、(2)总体校准(预测值和观察值的平均一致性)、(3)校准斜率(预测值和观察值在预测值范围内的一致性)和(4)观察值的方差。针对每个标准推导出封闭形式的样本量解决方案,这要求用户指定模型性能的预期值(特别是R-2)和外部验证数据集中的结果方差。明智的出发点是以模型开发研究的值为基础,如从出版物或研究作者那里获得的值。满足所有四个标准所需的最大样本量是外部验证数据集中所需的推荐最小样本量。当有固定样本大小的现有数据集可用时,这些计算也可以用于估计预期精度,以帮助衡量它是否足够。我们在一个预测儿童脱脂体重的案例研究中说明了所提出的方法。
Clinical prediction models provide individualized outcome predictions to inform patient counseling and clinical decision making. External validation is the process of examining a prediction model's performance in data independent to that used for model development. Current external validation studies often suffer from small sample sizes, and subsequently imprecise estimates of a model's predictive performance. To address this, we propose how to determine the minimum sample size needed for external validation of a clinical prediction model with a continuous outcome. Four criteria are proposed, that target precise estimates of (i) R-2 (the proportion of variance explained), (ii) calibration-in-the-large (agreement between predicted and observed outcome values on average), (iii) calibration slope (agreement between predicted and observed values across the range of predicted values), and (iv) the variance of observed outcome values. Closed-form sample size solutions are derived for each criterion, which require the user to specify anticipated values of the model's performance (in particular R-2) and the outcome variance in the external validation dataset. A sensible starting point is to base values on those for the model development study, as obtained from the publication or study authors. The largest sample size required to meet all four criteria is the recommended minimum sample size needed in the external validation dataset. The calculations can also be applied to estimate expected precision when an existing dataset with a fixed sample size is available, to help gauge if it is adequate. We illustrate the proposed methods on a case-study predicting fat-free mass in children.