On the choice and influence of the number of boosting steps for high-dimensional linear Cox-models

On the choice and influence of the number of boosting steps for high-dimensional linear Cox-models
复制标题

高维线性Cox模型boosting步数的选择及其影响

DOI:
--
复制
发表时间:
2016
期刊:
Computational statistics (Zeitschrift)
影响因子:
--
通讯作者:
Riccardo de Bin
Riccardo de Bin
中科院分区:
--
文献类型:
--
作者:
H. Seibold;C. Bernau;A. Boulesteix;Riccardo de Bin

文献摘要

参考文献

被引文献

相似文献

在生物医学研究中,基于提升的回归方法在过去十年中受到了广泛关注。它们的内在变量选择过程和将回归系数的估计值缩小到0的能力使得这些技术适合于在高维数据(例如基因表达)的情况下拟合预测模型。然而,它们的预测性能高度依赖于特定的调整参数,特别是要执行的提升迭代的数量。这个关键参数通常通过交叉验证来选择。交叉验证过程可以高度依赖于完全随机的分量,即所考虑的折叠分区。我们实证研究这种随机性在多大程度上影响了提升技术的结果,在选定的预测因子和相关模型的预测能力方面。我们使用了与四种不同疾病相关的四个公开数据集。在这些研究中,目标是在有大量连续候选预测因子可用时预测生存终点。我们专注于两个众所周知的升压方法中实现的R-包CoxBoost和mboost,假设的比例风险假设的有效性和线性的预测的影响。我们表明,在选定的预测和预测能力的模型的变异性降低平均在几个重复的交叉验证的调整参数的选择。
In biomedical research, boosting-based regression approaches have gained much attention in the last decade. Their intrinsic variable selection procedure and ability to shrink the estimates of the regression coefficients toward 0 make these techniques appropriate to fit prediction models in the case of high-dimensional data, e.g. gene expressions. Their prediction performance, however, highly depends on specific tuning parameters, in particular on the number of boosting iterations to perform. This crucial parameter is usually selected via cross-validation. The cross-validation procedure may highly depend on a completely random component, namely the considered fold partition. We empirically study how much this randomness affects the results of the boosting techniques, in terms of selected predictors and prediction ability of the related models. We use four publicly available data sets related to four different diseases. In these studies, the goal is to predict survival end-points when a large number of continuous candidate predictors are available. We focus on two well known boosting approaches implemented in the R-packages CoxBoost and mboost, assuming the validity of the proportional hazards assumption and the linearity of the effects of the predictors. We show that the variability in selected predictors and prediction ability of the model is reduced by averaging over several repetitions of cross-validation in the selection of the tuning parameters.
DOI: 10.1001/jama.2011.593
发表时间: 2011-05-11
期刊: JAMA
影响因子: --
作者:
Hatzis C;Pusztai L;Valero V;Booser DJ;Esserman L;Lluch A;Vidaurre T;Holmes F;Souchon E;Wang H;Martin M;Cotrina J;Gomez H;Hubbard R;Chacón JI;Ferrer-Lozano J;Dyer R;Buxton M;Gong Y;Wu Y;Ibrahim N;Andreopoulou E;Ueno NT;Hunt K;Yang W;Nazario A;DeMichele A;O'Shaughnessy J;Hortobagyi GN;Symmans WF
通讯作者: Symmans WF
DOI: 10.1182/blood-2008-02-134411
发表时间: 2008-11-15
期刊: BLOOD
影响因子: 20.3
作者:
Metzeler, Klaus H.;Hummel, Manuela;Buske, Christian
通讯作者: Buske, Christian
DOI: 10.1056/nejmoa012914
发表时间: 2002-06-20
影响因子: 158.5
作者:
Rosenwald, A;Wright, G;Staudt, LM
通讯作者: Staudt, LM
DOI: 10.18637/jss.v050.i11
发表时间: 2012-09
影响因子: 5.8
作者:
Mogensen UB;Ishwaran H;Gerds TA
通讯作者: Gerds TA