RISK PREDICTION FOR PROSTATE CANCER RECURRENCE THROUGH REGULARIZED ESTIMATION WITH SIMULTANEOUS ADJUSTMENT FOR NONLINEAR CLINICAL EFFECTS

RISK PREDICTION FOR PROSTATE CANCER RECURRENCE THROUGH REGULARIZED ESTIMATION WITH SIMULTANEOUS ADJUSTMENT FOR NONLINEAR CLINICAL EFFECTS
复制标题

DOI:
10.1214/11-aoas458
复制
发表时间:
2011-09-01
影响因子:
1.8
通讯作者:
Johnson, Brent A.
Johnson, Brent A.
中科院分区:
数学4区
文献类型:
--
作者:
Long, Qi;Chung, Matthias;Johnson, Brent A.

文献摘要

被引文献

相似文献

在生物医学研究中,使用高维数据(如基因表达数据)开发风险预测评分具有重大意义,这些数据适用于删失的临床终点。在存在明确的临床风险因素的情况下,研究者通常更喜欢也针对这些临床变量进行调整的程序。虽然加速失效时间(AFT)模型是分析截尾结果数据的有用工具,但它假设协变量对事件发生时间对数的影响是线性的,这在实践中往往是不现实的。我们建议通过部分线性AFT模型中的正则化秩估计来构建风险预测分数,其中高维数据(如基因表达数据)被线性建模,重要的临床变量使用惩罚回归样条非线性建模。我们通过仿真研究表明,我们的模型具有更好的操作特性相比,现有的几种模式。特别是,我们表明,当非线性临床效应被错误指定为线性时,对预测和特征选择有不可忽视的影响。这项工作的动机是最近的前列腺癌研究,其中研究人员收集基因表达数据沿着建立预后临床变量和主要终点是前列腺癌复发的时间。我们分析了前列腺癌数据,并基于删失数据的扩展c统计量评估了几种模型的预测性能,结果表明:(1)临床变量前列腺特异性抗原与前列腺癌复发之间的关系可能是非线性的,即复发时间随着PSA的增加而减少,当PSA> 11时,复发时间开始趋于平稳;(2)这种非线性效应的正确规范提高了预测和特征选择的性能;以及(3)基因表达数据的添加似乎没有进一步提高所得风险预测分数的性能。
In biomedical studies it is of substantial interest to develop risk prediction scores using high-dimensional data such as gene expression data for clinical endpoints that are subject to censoring. In the presence of well-established clinical risk factors, investigators often prefer a procedure that also adjusts for these clinical variables. While accelerated failure time (AFT) models are a useful tool for the analysis of censored outcome data, it assumes that covariate effects on the logarithm of time-to-event are linear, which is often unrealistic in practice. We propose to build risk prediction scores through regularized rank estimation in partly linear AFT models, where high-dimensional data such as gene expression data are modeled linearly and important clinical variables are modeled nonlinearly using penalized regression splines. We show through simulation studies that our model has better operating characteristics compared to several existing models. In particular, we show that there is a nonnegligible effect on prediction as well as feature selection when nonlinear clinical effects are misspecified as linear. This work is motivated by a recent prostate cancer study, where investigators collected gene expression data along with established prognostic clinical variables and the primary endpoint is time to prostate cancer recurrence. We analyzed the prostate cancer data and evaluated prediction performance of several models based on the extended c statistic for censored data, showing that (1) the relationship between the clinical variable, prostate specific antigen, and the prostate cancer recurrence is likely nonlinear, that is, the time to recurrence decreases as PSA increases and it starts to level off when PSA becomes greater than 11; (2) correct specification of this nonlinear effect improves performance in prediction and feature selection; and (3) addition of gene expression data does not seem to further improve the performance of the resultant risk prediction scores.