Minimum sample size calculations for external validation of a clinical prediction model with a time-to-event outcome

Minimum sample size calculations for external validation of a clinical prediction model with a time-to-event outcome
复制标题

DOI:
10.1002/sim.9275
复制
发表时间:
2021-12-16
影响因子:
2
通讯作者:
Snell, Kym I. E.
Snell, Kym I. E.
中科院分区:
医学3区
文献类型:
--
作者:
Riley, Richard D.;Collins, Gary S.;Snell, Kym I. E.

文献摘要

被引文献

相似文献

《医学统计学》之前的文章描述了如何计算具有连续和二元结果的预测模型的外部验证所需的样本量。最小样本量标准旨在确保准确估计模型预测性能的关键指标,包括校准、区分和净收益的指标。在这里,我们将样本量指导扩展到具有事件发生时间(生存)结果的预测模型,以涵盖包含审查的数据集中的外部验证。提出了一个基于模拟的框架,该框架计算了将校准斜率的特定可信区间宽度作为目标所需的样本量,以衡量在对数累积危险等级上预测风险(来自模型)和观察到的风险(使用伪观测来解释审查)之间的一致性。在这个框架中,还可以检查校准曲线、分辨率和净收益的精确估计。这一过程需要根据(I)模型线性预报器的分布和(Ii)事件和审查分布来假设验证总体。现有的信息可以提供这方面的信息;特别是,线性预测器分布可以使用模型开发文章中的C指数或Royston的D统计数据以及总体事件风险来近似。我们演示了如何使用该方法来计算验证复发静脉血栓栓塞症预测模型所需的样本量。理想情况下,样本量应该确保在整个预测风险范围内进行精确校准,但至少必须确保对临床决策重要的区域有足够的精确度。给出了STATA码和R码。
Previous articles in Statistics in Medicine describe how to calculate the sample size required for external validation of prediction models with continuous and binary outcomes. The minimum sample size criteria aim to ensure precise estimation of key measures of a model's predictive performance, including measures of calibration, discrimination, and net benefit. Here, we extend the sample size guidance to prediction models with a time-to-event (survival) outcome, to cover external validation in datasets containing censoring. A simulation-based framework is proposed, which calculates the sample size required to target a particular confidence interval width for the calibration slope measuring the agreement between predicted risks (from the model) and observed risks (derived using pseudo-observations to account for censoring) on the log cumulative hazard scale. Precise estimation of calibration curves, discrimination, and net-benefit can also be checked in this framework. The process requires assumptions about the validation population in terms of the (i) distribution of the model's linear predictor and (ii) event and censoring distributions. Existing information can inform this; in particular, the linear predictor distribution can be approximated using the C-index or Royston's D statistic from the model development article, together with the overall event risk. We demonstrate how the approach can be used to calculate the sample size required to validate a prediction model for recurrent venous thromboembolism. Ideally the sample size should ensure precise calibration across the entire range of predicted risks, but must at least ensure adequate precision in regions important for clinical decision-making. Stata and R code are provided.