Graphical calibration curves and the integrated calibration index (ICI) for survival models

Graphical calibration curves and the integrated calibration index (ICI) for survival models
复制标题

DOI:
10.1002/sim.8570
复制
发表时间:
2020-06-16
影响因子:
2
通讯作者:
van Klaveren, David
van Klaveren, David
中科院分区:
医学3区
文献类型:
--
作者:
Austin, Peter C.;Harrell, Frank E., Jr.;van Klaveren, David

文献摘要

被引文献

相似文献

在生存分析的背景下,校准是指在给定的持续时间内,预测的概率和观察到的事件发生率或频率之间的一致性。我们的目标是描述和评估图形化评估生存模型校准的方法。我们把重点放在风险回归模型和限制三次样条法以及COX比例风险模型上。我们还描述了对E50和E90的综合校准指数的修改。在这种情况下,这是预测生存概率和平滑生存频率之间的平均(分别为中位数或第90个百分位数)绝对差。我们进行了一系列蒙特卡罗模拟,以评估这些校准措施的性能,当基础模型已被正确指定时,以及在不同类型的模型错误指定下。我们通过比较Cox比例风险回归模型和随机生存森林模型的校准结果来说明校准曲线和三种校准指标在预测心力衰竭住院患者死亡率方面的作用。在正确指定的回归模型下,两种方法构建校准曲线的差异很小,尽管基于限制三次样条法的方法性能略好。相反,在错误指定的模型下,使用风险回归构建的平滑校准曲线往往更接近真正的校准曲线。校准曲线和这些数字校准指标的使用允许对竞争生存模型的校准进行全面比较。
In the context of survival analysis, calibration refers to the agreement between predicted probabilities and observed event rates or frequencies of the outcome within a given duration of time. We aimed to describe and evaluate methods for graphically assessing the calibration of survival models. We focus on hazard regression models and restricted cubic splines in conjunction with a Cox proportional hazards model. We also describe modifications of the Integrated Calibration Index, of E50 and of E90. In this context, this is the average (respectively, median or 90th percentile) absolute difference between predicted survival probabilities and smoothed survival frequencies. We conducted a series of Monte Carlo simulations to evaluate the performance of these calibration measures when the underlying model has been correctly specified and under different types of model mis-specification. We illustrate the utility of calibration curves and the three calibration metrics by using them to compare the calibration of a Cox proportional hazards regression model with that of a random survival forest for predicting mortality in patients hospitalized with heart failure. Under a correctly specified regression model, differences between the two methods for constructing calibration curves were minimal, although the performance of the method based on restricted cubic splines tended to be slightly better. In contrast, under a mis-specified model, the smoothed calibration curved constructed using hazard regression tended to be closer to the true calibration curve. The use of calibration curves and of these numeric calibration metrics permits for a comprehensive comparison of the calibration of competing survival models.