Assessment of performance of survival prediction models for cancer prognosis.

Assessment of performance of survival prediction models for cancer prognosis.
复制标题

DOI:
10.1186/1471-2288-12-102
复制
发表时间:
2012-07-23
影响因子:
4
通讯作者:
Chen JJ
Chen JJ
中科院分区:
医学3区
文献类型:
--
作者:
Chen HC;Kodell RL;Cheng KF;Chen JJ

文献摘要

参考文献

被引文献

相似文献

癌症生存研究通常使用生存时间预测模型进行癌症预后分析。许多不同的性能指标用于确定每个患者的预测风险评分与实际生存时间之间的一致性,但这些指标有时会发生冲突。或者,有时根据生存时间阈值将患者分为两类,并应用二元分类器来预测每个患者的类别。尽管这种方法有一些缺点,但它确实提供了自然的性能度量,例如正负预测值,以支持明确的评估。我们比较生存时间预测和生存时间阈值方法来分析癌症生存研究。我们回顾并比较了这两种方法的常见性能指标。我们提出了新的随机化测试和交叉验证方法,以实现与生存时间预测方法一起使用的几个性能指标的明确统计推断。我们考虑了五种生存预测模型,包括一种临床模型,两种基因表达模型,以及两种临床和基因表达模型的组合模型。一个公开的乳腺癌数据集被用来比较使用五种预测模型的几个性能指标。1)对于部分预测模型,Cox比例风险模型拟合的风险比显著,但两组比较不显著,反之亦然。2)随机化检验和交叉验证与标准性能指标得到的p值基本一致。3)二元分类器高度依赖于风险组的定义;类分配的生存阈值稍有变化,预测结果就会大不相同。1)评价生存预测模型的不同性能指标可能对其判别能力得出不同的结论。2)使用高风险和低风险组比较的评估取决于所选择的风险评分阈值;所有可能阈值的p值图可以显示阈值选择的敏感性。3) Somers等级相关显著性的随机化检验可用于进一步评价预测模型的性能。4)生存预测模型的交叉验证能力随着训练集和测试集的不平衡而降低。
Cancer survival studies are commonly analyzed using survival-time prediction models for cancer prognosis. A number of different performance metrics are used to ascertain the concordance between the predicted risk score of each patient and the actual survival time, but these metrics can sometimes conflict. Alternatively, patients are sometimes divided into two classes according to a survival-time threshold, and binary classifiers are applied to predict each patient’s class. Although this approach has several drawbacks, it does provide natural performance metrics such as positive and negative predictive values to enable unambiguous assessments. We compare the survival-time prediction and survival-time threshold approaches to analyzing cancer survival studies. We review and compare common performance metrics for the two approaches. We present new randomization tests and cross-validation methods to enable unambiguous statistical inferences for several performance metrics used with the survival-time prediction approach. We consider five survival prediction models consisting of one clinical model, two gene expression models, and two models from combinations of clinical and gene expression models. A public breast cancer dataset was used to compare several performance metrics using five prediction models. 1) For some prediction models, the hazard ratio from fitting a Cox proportional hazards model was significant, but the two-group comparison was insignificant, and vice versa. 2) The randomization test and cross-validation were generally consistent with the p-values obtained from the standard performance metrics. 3) Binary classifiers highly depended on how the risk groups were defined; a slight change of the survival threshold for assignment of classes led to very different prediction results. 1) Different performance metrics for evaluation of a survival prediction model may give different conclusions in its discriminatory ability. 2) Evaluation using a high-risk versus low-risk group comparison depends on the selected risk-score threshold; a plot of p-values from all possible thresholds can show the sensitivity of the threshold selection. 3) A randomization test of the significance of Somers’ rank correlation can be used for further evaluation of performance of a prediction model. 4) The cross-validated power of survival prediction models decreases as the training and test sets become less balanced.
DOI: 10.1038/35000501
发表时间: 2000-02-03
期刊: NATURE
影响因子: 64.8
作者:
Alizadeh, AA;Eisen, MB;Staudt, LM
通讯作者: Staudt, LM
DOI: 10.1371/journal.pbio.0020108
发表时间: 2004-04
期刊: PLoS biology
影响因子: 9.8
作者:
Bair E;Tibshirani R
通讯作者: Tibshirani R
DOI: 10.1016/j.ejca.2006.11.018
发表时间: 2007-03-01
影响因子: 8.4
作者:
Dunkler, Daniela;Michiels, Stefan;Schemper, Michael
通讯作者: Schemper, Michael
DOI: 10.1200/jco.2003.01.240
发表时间: 2003-10-01
影响因子: 45.3
作者:
Kattan, MW;Karpeh, MS;Brennan, MF
通讯作者: Brennan, MF