The sparsity and bias of the lasso selection in high-dimensional linear regression

The sparsity and bias of the lasso selection in high-dimensional linear regression
复制标题

DOI:
10.1214/07-aos520
复制
发表时间:
2008-08-01
影响因子:
4.5
通讯作者:
Huang, Jian
Huang, Jian
中科院分区:
数学1区
文献类型:
--
作者:
Zhang, Cun-Hui;Huang, Jian

文献摘要

被引文献

相似文献

Meinshausen和Buhlmann [Ann. Statistist. 34(2006)1436-1462]表明,对于高斯图形模型中的邻域选择,在邻域稳定性条件下,LASSO是一致的,即使当变量的数量比样本大小更高阶时。Zhao和Yu [(2006)J. Machine Learning Research 7 2541-2567]将线性回归上下文中的邻域稳定性条件形式化为强不可表示条件。该论文表明,在这种条件下,LASSO选择的正是非零回归系数的集合,只要这些系数以一定的速率远离零。本文假设理想模型外的回归系数很小,但不一定为零。在稀疏的Riesz条件下的设计变量的相关性,我们证明了LASSO选择一个模型的正确顺序的维度,控制所选模型的偏差在一个水平上的小回归系数和阈值偏差的贡献所确定的,并选择所有系数的更大的顺序比所选模型的偏差。此外,由于这种速率一致性的LASSO模型选择,它被证明是误差平方和的平均响应和l(alpha)-损失的回归系数收敛在最佳可能的速率在给定的条件下。我们的结果的一个有趣的方面是,变量的数量的对数可以是相同的顺序作为某些随机相关设计的样本容量。
Meinshausen and Buhlmann [Ann. Statist. 34 (2006) 1436-1462] showed that, for neighborhood selection in Gaussian graphical models, under a neighborhood stability condition, the LASSO is consistent, even when the number of variables is of greater order than the sample size. Zhao and Yu [(2006) J. Machine Learning Research 7 2541-2567] formalized the neighborhood stability condition in the context of linear regression as a strong irrepresentable condition. That paper showed that under this condition, the LASSO selects exactly the set of nonzero regression coefficients, provided that these coefficients are bounded away from zero at a certain rate. In this paper, the regression coefficients outside an ideal model are assumed to be small, but not necessarily zero. Under a sparse Riesz condition on the correlation of design variables, we prove that the LASSO selects a model of the correct order of dimensionality, controls the bias of the selected model at a level determined by the contributions of small regression coefficients and threshold bias, and selects all coefficients of greater order than the bias of the selected model. Moreover, as a consequence of this rate consistency of the LASSO in model selection, it is proved that the sum of error squares for the mean response and the l(alpha)-loss for the regression coefficients converge at the best possible rates under the given conditions. An interesting aspect of our results is that the logarithm of the number of variables can be of the same order as the sample size for certain random dependent designs.