Nonpenalized variable selection in high-dimensional linear model settings via generalized fiducial inference

Nonpenalized variable selection in high-dimensional linear model settings via generalized fiducial inference
复制标题

DOI:
10.1214/18-aos1733
复制
发表时间:
2017-02
期刊:
The Annals of Statistics
影响因子:
--
通讯作者:
Jonathan P. Williams;Jan Hannig
Jonathan P. Williams;Jan Hannig
中科院分区:
其他
文献类型:
--
作者:
Jonathan P. Williams;Jan Hannig

文献摘要

被引文献

相似文献

变量选择和参数估计的标准惩罚方法依赖于系数估计值的大小来决定最终模型中包含哪些变量。然而,当设计矩阵共线时,系数估计值是不可靠的。为了克服这一挑战,一个全新的角度变量选择的广义基准推理框架内。这个新的过程是能够有效地考虑线性依赖关系的协变量的子集之间的高维设置,其中$p$可以增长几乎指数在$n$,以及在经典的设置,其中$p \le n$。它表明,该程序非常自然地分配小概率的协变量的子集,其中包括冗余的方式显式L_{0}$最小化。此外,与一个典型的稀疏性假设,它表明,所提出的方法是一致的,在这个意义上,真正的稀疏子集的协变量的概率收敛到1作为$n \到\infty$,或作为$n \到\infty$和$p \到\infty$。需要非常合理的条件,并且对协变量的可能子集的类别几乎没有限制,以实现这种一致性结果。
Standard penalized methods of variable selection and parameter estimation rely on the magnitude of coefficient estimates to decide which variables to include in the final model. However, coefficient estimates are unreliable when the design matrix is collinear. To overcome this challenge an entirely new perspective on variable selection is presented within a generalized fiducial inference framework. This new procedure is able to effectively account for linear dependencies among subsets of covariates in a high-dimensional setting where $p$ can grow almost exponentially in $n$, as well as in the classical setting where $p \le n$. It is shown that the procedure very naturally assigns small probabilities to subsets of covariates which include redundancies by way of explicit $L_{0}$ minimization. Furthermore, with a typical sparsity assumption, it is shown that the proposed method is consistent in the sense that the probability of the true sparse subset of covariates converges in probability to 1 as $n \to \infty$, or as $n \to \infty$ and $p \to \infty$. Very reasonable conditions are needed, and little restriction is placed on the class of possible subsets of covariates to achieve this consistency result.