On the use of Harrell's C for clinical risk prediction via random survival forests

On the use of Harrell's C for clinical risk prediction via random survival forests
复制标题

DOI:
10.1016/j.eswa.2016.07.018
复制
发表时间:
2016-11-30
影响因子:
8.5
通讯作者:
Ziegler, Andreas
Ziegler, Andreas
中科院分区:
计算机科学1区
文献类型:
--
作者:
Schmid, Matthias;Wright, Marvin N.;Ziegler, Andreas

文献摘要

被引文献

相似文献

随机生存森林(RSF)是生物医学研究中右删失结果风险预测的一种有效方法。RSF使用对数秩分割标准来形成生存树的集合。评估RSF模型预测准确性的最常用方法是Harrell生存数据一致性指数(“C指数”)。从概念上讲,这种策略意味着RSF中的分割标准与感兴趣的评估标准不同。这种差异可以通过使用Harrell's C进行节点分裂和评估来克服。我们比较了两个分裂标准之间的差异分析和模拟研究方面的偏好更不平衡的分裂,称为端切偏好(ECP)。具体来说,我们表明,对数秩统计量有一个更强的ECP相比,C指数。在模拟研究和两个医疗数据集的帮助下,我们证明了RSF预测的准确性,如测量Harrell的C,可以提高,如果对数秩统计被取代的C指数节点分裂。这在截尾率或信息连续预测变量的分数很高的情况下尤其如此。相反,对数秩分裂在噪声场景中是优选的。基于C的和对数秩分裂都在R包中实现。我们推荐Harrell's C作为小规模临床研究的分割标准,并推荐对数秩分割标准用于大规模组学研究。(C)2016由Elsevier Ltd.出版
Random survival forests (RSF) are a powerful method for risk prediction of right-censored outcomes in biomedical research. RSF use the log-rank split criterion to form an ensemble of survival trees. The most common approach to evaluate the prediction accuracy of a RSF model is Harrell's concordance index for survival data ('C index'). Conceptually, this strategy implies that the split criterion in RSF is different from the evaluation criterion of interest. This discrepancy can be overcome by using Harrell's C for both node splitting and evaluation. We compare the difference between the two split criteria analytically and in simulation studies with respect to the preference of more unbalanced splits, termed end-cut preference (ECP). Specifically, we show that the log-rank statistic has a stronger ECP compared to the C index. In simulation studies and with the help of two medical data sets we demonstrate that the accuracy of RSF predictions, as measured by Harrell's C, can be improved if the log-rank statistic is replaced by the C index for node splitting. This is especially true in situations where the censoring rate or the fraction of informative continuous predictor variables is high. Conversely, log-rank splitting is preferable in noisy scenarios. Both C-based and log-rank splitting are implemented in the R package ranger. We recommend Harrell's C as split criterion for use in smaller scale clinical studies and the log-rank split criterion for use in large-scale 'omics' studies. (C) 2016 Published by Elsevier Ltd.