Meta-analysis of prediction model performance across multiple studies: Which scale helps ensure between-study normality for the C-statistic and calibration measures?

Meta-analysis of prediction model performance across multiple studies: Which scale helps ensure between-study normality for the C-statistic and calibration measures?
复制标题

对多项研究的预测模型性能的荟萃分析:哪个量表有助于确保C统计和校准度量的研究中的正态性?

DOI:
10.1177/0962280217705678
复制
发表时间:
2018-11
影响因子:
2.3
通讯作者:
Riley RD
Riley RD
中科院分区:
医学3区
文献类型:
--
作者:
Snell KI;Ensor J;Debray TP;Moons KG;Riley RD

文献摘要

参考文献

被引文献

相似文献

如果单个参与者数据可从多个研究或聚类中获得,则预测模型可以多次进行外部验证。这允许在不同的设置下检查模型的识别和校准性能。随机效应荟萃分析可以用来量化整体(平均)性能和性能的异质性。这通常假设研究中的“真实”性能呈正态分布。我们进行了模拟研究,以检查这种正态性假设的各种性能指标相关的逻辑回归预测模型。我们模拟了多项研究的数据,这些研究在基线风险或预测效应方面具有不同程度的变异性,然后评估了C统计量、校准斜率、大规模校准和E/O统计量中研究间分布的形状及其可能的转换。我们发现,研究间正态分布通常对于校准斜率和大规模校准是合理的;然而,C统计量和E/O的分布通常在研究间偏斜,特别是在预测效应具有较大变异性的环境中。当对C-统计量使用logit转换和对E/O使用log转换时,正态性得到了极大的改善,因此我们建议将这些量表用于荟萃分析。一个说明性的例子是使用QRISK 2在25个一般做法的性能的随机效应荟萃分析。
If individual participant data are available from multiple studies or clusters, then a prediction model can be externally validated multiple times. This allows the model’s discrimination and calibration performance to be examined across different settings. Random-effects meta-analysis can then be used to quantify overall (average) performance and heterogeneity in performance. This typically assumes a normal distribution of ‘true’ performance across studies. We conducted a simulation study to examine this normality assumption for various performance measures relating to a logistic regression prediction model. We simulated data across multiple studies with varying degrees of variability in baseline risk or predictor effects and then evaluated the shape of the between-study distribution in the C-statistic, calibration slope, calibration-in-the-large, and E/O statistic, and possible transformations thereof. We found that a normal between-study distribution was usually reasonable for the calibration slope and calibration-in-the-large; however, the distributions of the C-statistic and E/O were often skewed across studies, particularly in settings with large variability in the predictor effects. Normality was vastly improved when using the logit transformation for the C-statistic and the log transformation for E/O, and therefore we recommend these scales to be used for meta-analysis. An illustrated example is given using a random-effects meta-analysis of the performance of QRISK2 across 25 general practices.
DOI: 10.1093/aje/kwt298
发表时间: 2014-03-01
影响因子: 5
作者:
Pennells L;Kaptoge S;White IR;Thompson SG;Wood AM;Emerging Risk Factors Collaboration
通讯作者: Emerging Risk Factors Collaboration
DOI: 10.1186/1471-2288-12-82
发表时间: 2012-06-20
影响因子: 4
作者:
Austin PC;Steyerberg EW
通讯作者: Steyerberg EW
DOI: 10.1016/j.jclinepi.2016.05.007
发表时间: 2016-11-01
影响因子: 7.2
作者:
Austin, Peter C.;van Klaveren, David;Steyerberg, Ewout W.
通讯作者: Steyerberg, Ewout W.
DOI: 10.7326/0003-4819-130-6-199903160-00016
发表时间: 1999-03-16
影响因子: 39.2
作者:
Justice, AC;Covinsky, KE;Berlin, JA
通讯作者: Berlin, JA
DOI: 10.1136/bmj.d549
发表时间: 2011-02-10
影响因子: 105.7
作者:
Riley, Richard D.;Higgins, Julian P. T.;Deeks, Jonathan J.
通讯作者: Deeks, Jonathan J.