Boosting for high-dimensional linear models

Boosting for high-dimensional linear models
复制标题

DOI:
10.1214/009053606000000092
复制
发表时间:
2006-04-01
影响因子:
4.5
通讯作者:
Buhlmann, Peter
Buhlmann, Peter
中科院分区:
数学1区
文献类型:
--
作者:
Buhlmann, Peter

文献摘要

被引文献

相似文献

我们证明了具有平方误差损失的Boosting,L(2)Boosting,对于非常高维的线性模型是一致的,其中允许预测变量的数量基本上以O(exp(样本大小))的速度增长,假设真实的底层回归函数在回归系数的l(1)范数方面是稀疏的。在信号处理的语言中,这意味着如果底层信号在l(1)范数方面是稀疏的,则使用强过完备字典进行去噪的一致性。我们还在这里提出了一个基于AIC的调整方法,即选择提升迭代次数。这使得L(2)Boosting在计算上具有吸引力,因为它不需要像目前为止常用的那样多次运行算法进行交叉验证。我们证明了L(2)提升的模拟数据,特别是在预测维度是大的样本量相比,和一个困难的肿瘤分类问题与基因表达微阵列数据。
We prove that boosting with the squared error loss, L(2)Boosting, is consistent for very high-dimensional linear models, where the number of predictor variables is allowed to grow essentially as fast as O(exp(sample size)), assuming that the true underlying regression function is sparse in terms of the l(1)-norm of the regression coefficients. In the language of signal processing, this means consistency for de-noising using a strongly overcomplete dictionary if the underlying signal is sparse in terms of the l(1)-norm. We also propose here an AIC-based method for tuning, namely for choosing the number of boosting iterations. This makes L(2)Boosting computationally attractive since it is not required to run the algorithm multiple times for cross-validation as commonly used so far. We demonstrate L(2)Boosting for simulated data, in particular where the predictor dimension is large in comparison to sample size, and for a difficult tumor-classification problem with gene expression microarray data.