The Lasso with general Gaussian designs with applications to hypothesis testing

The Lasso with general Gaussian designs with applications to hypothesis testing
复制标题

DOI:
10.1214/23-aos2327
复制
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Michael Celentano;A. Montanari;Yuting Wei
Michael Celentano;A. Montanari;Yuting Wei
中科院分区:
其他
文献类型:
--
作者:
Michael Celentano;A. Montanari;Yuting Wei

文献摘要

被引文献

相似文献

Lasso是一种用于高维回归的方法,现在通常用于协变量$p$的数量与观测值$n$的数量相同或更大。经典的渐近正态性理论不适用于该模型,这有两个根本原因:$(1)$正则化风险是非光滑的; $(2)$估计量$\bf \widehat{\theta}$与真实参数向量$\bf \theta^\星星$之间的距离不可忽略。因此,作为渐近正态性的传统基础的标准微扰参数失败了。另一方面,Lasso估计量可以在n和p都很大,而n/p是一阶的情况下精确地表征。这种特征首先在标准高斯设计的情况下得到,随后推广到其他高维估计程序。在这里,我们将相同的特征扩展到具有非奇异协方差结构的高斯相关设计。这一特征用一个更简单的"固定设计“模型来表示。我们建立了两个模型中的各种数量的分布之间的距离的非渐近界,这两个模型中的信号$\bf \theta^\星星$在一个合适的稀疏类,和正则化参数的值一致。作为应用,我们研究了去偏Lasso的分布,并证明了自由度校正对于计算有效的置信区间是必要的。
The Lasso is a method for high-dimensional regression, which is now commonly used when the number of covariates $p$ is of the same order or larger than the number of observations $n$. Classical asymptotic normality theory is not applicable for this model due to two fundamental reasons: $(1)$ The regularized risk is non-smooth; $(2)$ The distance between the estimator $\bf \widehat{\theta}$ and the true parameters vector $\bf \theta^\star$ cannot be neglected. As a consequence, standard perturbative arguments that are the traditional basis for asymptotic normality fail. On the other hand, the Lasso estimator can be precisely characterized in the regime in which both $n$ and $p$ are large, while $n/p$ is of order one. This characterization was first obtained in the case of standard Gaussian designs, and subsequently generalized to other high-dimensional estimation procedures. Here we extend the same characterization to Gaussian correlated designs with non-singular covariance structure. This characterization is expressed in terms of a simpler ``fixed design'' model. We establish non-asymptotic bounds on the distance between distributions of various quantities in the two models, which hold uniformly over signals $\bf \theta^\star$ in a suitable sparsity class, and values of the regularization parameter. As applications, we study the distribution of the debiased Lasso, and show that a degrees-of-freedom correction is necessary for computing valid confidence intervals.