Least squares after model selection in high-dimensional sparse models

Least squares after model selection in high-dimensional sparse models
复制标题

DOI:
10.3150/11-bej410
复制
发表时间:
2013-05-01
期刊:
影响因子:
1.5
通讯作者:
Chernozhukov, Victor
Chernozhukov, Victor
中科院分区:
数学2区
文献类型:
--
作者:
Belloni, Alexandre;Chernozhukov, Victor

文献摘要

被引文献

相似文献

在这篇文章中,我们研究后模型选择估计,适用于普通最小二乘(OLS)的第一步惩罚估计,典型的Lasso选择的模型。众所周知,Lasso可以以接近预言的速度估计非参数回归函数,因此很难改进。我们表明,OLS后Lasso估计执行至少以及Lasso的收敛速度,并具有较小的偏差的优势。值得注意的是,即使基于Lasso的模型选择“失败”,也会出现这种性能,因为它丢失了“真正”回归模型的一些组成部分。所谓“真”模型,我们指的是对预言机选择的非参数回归函数的最佳s维近似。此外,如果基于Lasso的模型选择正确地包括“真”模型的所有分量作为子集,并且还实现了足够的稀疏性,则OLS post-Lasso估计量可以在收敛速度更快的意义上严格优于Lasso。在极端情况下,当Lasso完美地选择“真”模型时,OLS post-Lasso估计量成为oracle估计量。在我们的分析中,一个重要的成分是一个新的稀疏约束的维度选择的模型的Lasso,这保证了这个维度是在最多的相同的顺序作为“真正的”模型的维度。我们的速率结果是非渐近的,并在参数和非参数模型。此外,我们的分析不仅限于在第一步中作为选择器的Lasso估计器,而且还适用于任何其他估计器,例如,各种形式的阈值Lasso,具有良好的速率和良好的稀疏性。我们的分析涵盖了传统的阈值和一个新的实用,数据驱动的阈值方案,诱导额外的稀疏性保持一定的拟合优度。后一种方案具有类似于Lasso或OLS post-Lasso的理论保证,但它在各种实验中主导这些程序以及传统阈值。
In this article we study post-model selection estimators that apply ordinary least squares (OLS) to the model selected by first-step penalized estimators, typically Lasso. It is well known that Lasso can estimate the nonparametric regression function at nearly the oracle rate, and is thus hard to improve upon. We show that the OLS post-Lasso estimator performs at least as well as Lasso in terms of the rate of convergence, and has the advantage of a smaller bias. Remarkably, this performance occurs even if the Lasso-based model selection "fails" in the sense of missing some components of the "true" regression model. By the "true" model, we mean the best s-dimensional approximation to the nonparametric regression function chosen by the oracle. Furthermore, OLS post-Lasso estimator can perform strictly better than Lasso, in the sense of a strictly faster rate of convergence, if the Lasso-based model selection correctly includes all components of the "true" model as a subset and also achieves sufficient sparsity. In the extreme case, when Lasso perfectly selects the "true" model, the OLS post-Lasso estimator becomes the oracle estimator. An important ingredient in our analysis is a new sparsity bound on the dimension of the model selected by Lasso, which guarantees that this dimension is at most of the same order as the dimension of the "true" model. Our rate results are nonasymptotic and hold in both parametric and nonparametric models. Moreover, our analysis is not limited to the Lasso estimator acting as a selector in the first step, but also applies to any other estimator, for example, various forms of thresholded Lasso, with good rates and good sparsity properties. Our analysis covers both traditional thresholding and a new practical, data-driven thresholding scheme that induces additional sparsity subject to maintaining a certain goodness of fit. The latter scheme has theoretical guarantees similar to those of Lasso or OLS post-Lasso, but it dominates those procedures as well as traditional thresholding in a wide variety of experiments.