Generalizability of Cardiovascular Disease Clinical Prediction Models: 158 Independent External Validations of 104 Unique Models.

Generalizability of Cardiovascular Disease Clinical Prediction Models: 158 Independent External Validations of 104 Unique Models.
复制标题

DOI:
10.1161/circoutcomes.121.008487
复制
发表时间:
2022-04
影响因子:
6.9
通讯作者:
Kent, David M.
Kent, David M.
中科院分区:
医学1区
文献类型:
--
作者:
Gulati, Gaurav;Upshaw, Jenica;Wessler, Benjamin S.;Brazil, Riley J.;Nelson, Jason;van Klaveren, David;Lundquist, Christine M.;Park, Jinny G.;McGinnes, Hannah;Steyerberg, Ewout W.;Van Calster, Ben;Kent, David M.

文献摘要

被引文献

相似文献

虽然临床预测模型(CPM)越来越普遍地用于指导患者护理,但对这些CPM在新患者队列中的性能和临床效用了解甚少。我们对心血管疾病3个领域(一级预防、急性冠状动脉综合征和心力衰竭)的104种独特CPM进行了158项外部验证。在公开的临床试验队列中进行验证,并使用区分度、校准和净效益指标评估模型性能。为了探索模型性能不佳的潜在原因,基于相关性对CPM-临床试验队列对进行分层,相关性是一组特定领域的特征,用于对推导和验证患者人群的相似性进行定性分级。我们还检查了基于模型的C-统计量,以评估歧视的变化是否是因为衍生和验证样本之间的病例组合差异。还评估了模型更新对模型性能的影响。模型推导(0.76 [四分位距0.73-0.78])和验证(0.64 [四分位距0.60-0.67],P<0.001)之间的区分度显著降低,但这种降低的大约一半是因为验证样本中的病例组合较窄。与远亲试验队列相比,CPM在相关试验中具有更好的辨别力。相关试验队列中的校准斜率(0.77 [四分位距,0.59-0.90])也显著高于远亲队列(0.59 [四分位距0.43-0.73],P=0.001)。当考虑结果发生率的一半和两倍之间的可能决策阈值的全部范围时,91%的模型在某个阈值处具有损害风险(净受益低于默认策略);该风险可以通过更新模型截距、校准斜率或完全重新估计来大幅降低。当将心血管疾病CPM应用于新患者人群时,模型性能会显着下降,从而导致巨大的伤害风险。模型更新可以减轻这些风险。使用CPM指导临床决策时应谨慎。
While clinical prediction models (CPMs) are used increasingly commonly to guide patient care, the performance and clinical utility of these CPMs in new patient cohorts is poorly understood. We performed 158 external validations of 104 unique CPMs across 3 domains of cardiovascular disease (primary prevention, acute coronary syndrome, and heart failure). Validations were performed in publicly available clinical trial cohorts and model performance was assessed using measures of discrimination, calibration, and net benefit. To explore potential reasons for poor model performance, CPM-clinical trial cohort pairs were stratified based on relatedness, a domain-specific set of characteristics to qualitatively grade the similarity of derivation and validation patient populations. We also examined the model-based C-statistic to assess whether changes in discrimination were because of differences in case-mix between the derivation and validation samples. The impact of model updating on model performance was also assessed. Discrimination decreased significantly between model derivation (0.76 [interquartile range 0.73–0.78]) and validation (0.64 [interquartile range 0.60–0.67], P<0.001), but approximately half of this decrease was because of narrower case-mix in the validation samples. CPMs had better discrimination when tested in related compared with distantly related trial cohorts. Calibration slope was also significantly higher in related trial cohorts (0.77 [interquartile range, 0.59–0.90]) than distantly related cohorts (0.59 [interquartile range 0.43–0.73], P=0.001). When considering the full range of possible decision thresholds between half and twice the outcome incidence, 91% of models had a risk of harm (net benefit below default strategy) at some threshold; this risk could be reduced substantially via updating model intercept, calibration slope, or complete re-estimation. There are significant decreases in model performance when applying cardiovascular disease CPMs to new patient populations, resulting in substantial risk of harm. Model updating can mitigate these risks. Care should be taken when using CPMs to guide clinical decision-making.