Minimum sample size for external validation of a clinical prediction model with a binary outcome

Minimum sample size for external validation of a clinical prediction model with a binary outcome
复制标题

DOI:
10.1002/sim.9025
复制
发表时间:
2021-05-24
影响因子:
2
通讯作者:
Snell, Kym I. E.
Snell, Kym I. E.
中科院分区:
医学3区
文献类型:
--
作者:
Riley, Richard D.;Debray, Thomas P. A.;Snell, Kym I. E.

文献摘要

被引文献

相似文献

在预测模型研究中,需要外部验证来使用独立于模型开发的数据来检验现有模型的性能。目前的外部验证研究往往存在样本量小,因此预测性能估计不准确的问题。为了解决这一问题,我们建议如何确定对二元结果的预测模型进行新的外部验证研究所需的最小样本量。我们的计算旨在精确地估计校准(观察/预期和校准斜率)、辨别度(C-统计量)和临床效用(净收益)。对于每个度量,我们提出了计算所需最小样本量的闭合形式和迭代解。这些要求规定:(I)每个感兴趣估计的目标SE(可信区间宽度),(Ii)验证总体中的预期结果事件比例,(Iii)预测模型在验证总体中线性预测器值的预期(误)校准和方差,以及(Iv)临床决策的潜在风险阈值。这些计算还可用于通知现有(已收集的)数据集的样本大小是否足以进行外部验证。我们举例说明了我们的建议,即外部验证机械性心脏瓣膜衰竭的预测模型,其预期结果事件比例为0.018。计算表明,至少需要9835名参与者(177个事件)来准确估计校准和区分措施,这一数字由校准斜率标准决定,我们预计通常会出现这种情况。此外,6443名参与者(116项活动)需要在8%的风险阈值下准确估计净收益。提供了软件代码。
In prediction model research, external validation is needed to examine an existing model's performance using data independent to that for model development. Current external validation studies often suffer from small sample sizes and consequently imprecise predictive performance estimates. To address this, we propose how to determine the minimum sample size needed for a new external validation study of a prediction model for a binary outcome. Our calculations aim to precisely estimate calibration (Observed/Expected and calibration slope), discrimination (C-statistic), and clinical utility (net benefit). For each measure, we propose closed-form and iterative solutions for calculating the minimum sample size required. These require specifying: (i) target SEs (confidence interval widths) for each estimate of interest, (ii) the anticipated outcome event proportion in the validation population, (iii) the prediction model's anticipated (mis)calibration and variance of linear predictor values in the validation population, and (iv) potential risk thresholds for clinical decision-making. The calculations can also be used to inform whether the sample size of an existing (already collected) dataset is adequate for external validation. We illustrate our proposal for external validation of a prediction model for mechanical heart valve failure with an expected outcome event proportion of 0.018. Calculations suggest at least 9835 participants (177 events) are required to precisely estimate the calibration and discrimination measures, with this number driven by the calibration slope criterion, which we anticipate will often be the case. Also, 6443 participants (116 events) are required to precisely estimate net benefit at a risk threshold of 8%. Software code is provided.