Efficient test-based variable selection for high-dimensional linear models

Efficient test-based variable selection for high-dimensional linear models
复制标题

高维线性模型基于测试的高效变量选择

DOI:
10.1016/j.jmva.2018.01.003
复制
发表时间:
2018
影响因子:
1.6
通讯作者:
Liu, Yufeng
Liu, Yufeng
中科院分区:
数学2区
文献类型:
--
作者:
Gong, Siliang;Zhang, Kai;Liu, Yufeng

文献摘要

参考文献

相似文献

变量选择在高维数据分析中起着基础性的作用。近年来,已经开发了各种方法用于变量选择。众所周知的例子是前向逐步回归(FSR)和最小角度回归(LARS)等。这些方法通常将变量逐个添加到模型中。对于这样的选择过程,找到一个控制模型复杂性的停止准则是至关重要的。为此,最常用的技术之一是交叉验证(CV),尽管它很受欢迎,但它有两个主要缺点:昂贵的计算成本和缺乏统计解释。为了克服这些缺点,我们引入了一个灵活和高效的基于测试的变量选择方法,可以纳入任何顺序选择程序。该测试是对剩余非活动变量中的总体信号的测试,其基于非活动变量与给定活动变量的响应之间的最大绝对偏相关。我们发展的渐近零分布的检验统计量的维数趋于无穷大的样本容量一致。我们还表明,测试是一致的。使用此检验,在选择的每个步骤中,当且仅当p值低于某个预定义水平时,才包含一个新变量。数值研究表明,与CV相比,该方法在变量选择准确性和计算复杂度方面具有非常有竞争力的性能。
Variable selection plays a fundamental role in high-dimensional data analysis. Various methods have been developed for variable selection in recent years. Well-known examples are forward stepwise regression (FSR) and least angle regression (LARS), among others. These methods typically add variables into the model one by one. For such selection procedures, it is crucial to find a stopping criterion that controls model complexity. One of the most commonly used techniques to this end is cross-validation (CV) which, in spite of its popularity, has two major drawbacks: expensive computational cost and lack of statistical interpretation. To overcome these drawbacks, we introduce a flexible and efficient test-based variable selection approach that can be incorporated into any sequential selection procedure. The test, which is on the overall signal in the remaining inactive variables, is based on the maximal absolute partial correlation between the inactive variables and the response given active variables. We develop the asymptotic null distribution of the proposed test statistic as the dimension tends to infinity uniformly in the sample size. We also show that the test is consistent. With this test, at each step of the selection, a new variable is included if and only if the p-value is below some pre-defined level. Numerical studies show that the proposed method delivers very competitive performance in terms of variable selection accuracy and computational complexity compared to CV.
DOI: 10.1111/rssb.12048
发表时间: 2014-09-01
影响因子: 5.8
作者:
Aharoni, Ehud;Rosset, Saharon
通讯作者: Rosset, Saharon
DOI: 10.1214/13-aos1175
发表时间: 2014-04
影响因子: 4.5
作者:
Lockhart R;Taylor J;Tibshirani RJ;Tibshirani R
通讯作者: Tibshirani R
DOI: 10.1109/tit.2017.2700202
发表时间: 2017
影响因子: 2.5
作者:
Zhang, Kai
通讯作者: Zhang, Kai
DOI: 10.1093/nsr/nwt032
发表时间: 2014-06
影响因子: 20.6
作者:
Fan J;Han F;Liu H
通讯作者: Liu H