NEARLY UNBIASED VARIABLE SELECTION UNDER MINIMAX CONCAVE PENALTY

NEARLY UNBIASED VARIABLE SELECTION UNDER MINIMAX CONCAVE PENALTY
复制标题

DOI:
10.1214/09-aos729
复制
发表时间:
2010-04-01
影响因子:
4.5
通讯作者:
Zhang, Cun-Hui
Zhang, Cun-Hui
中科院分区:
数学1区
文献类型:
--
作者:
Zhang, Cun-Hui

文献摘要

被引文献

相似文献

提出了一种快速、连续、几乎无偏、准确的高维线性回归惩罚变量选择方法MC+。套索快速而连续,但有偏向。套索的偏向可能会妨碍一致的变量选择。子集选择是无偏见的,但计算代价很高。MC+有两个元素:极小极大凹惩罚(MCP)和惩罚线性无偏选择(PLUS)算法。在给定变量选择和无偏的某些阈值的情况下,MCP在最大程度上提供了稀疏区域中惩罚损失的凸性。加法在惩罚损失临界点的图的某一主分支上计算可能非凸的惩罚损失函数的多个精确局部极小值。它的输出是一条连续的分段线性路径,从无限罚金的原点到零罚金的最小二乘解。我们证明了在普遍惩罚水平下,MC+有很高的概率匹配未知数的符号,从而正确地选择,而不假设套索所要求的强不可表示条件。这种选择一致性适用于p>>n的情形,并且被证明在可能的多个局部极小值中恰好适用于MC+解。证明了MC+估计e球回归系数在概率上达到了一定的极小极大收敛速度。利用Sure方法得到了一般惩罚最小二乘估计的自由度和C-p型风险估计,包括LASSO估计和MC+估计,并证明了它们的无偏性。基于估计的自由度,我们提出了噪声水平的估计器,以便适当地选择惩罚水平。对于满秩次设计和一般次二次惩罚,我们给出了惩罚LSE连续的充要条件。仿真结果压倒性地支持了我们关于变量选择优性的主张,并证明了该方法的计算效率。
We propose MC+, a fast, continuous, nearly unbiased and accurate method of penalized variable selection in high-dimensional linear regression. The LASSO is fast and continuous, but biased. The bias of the LASSO may prevent consistent variable selection. Subset selection is unbiased but computationally costly. The MC+ has two elements: a minimax concave penalty (MCP) and a penalized linear unbiased selection (PLUS) algorithm. The MCP provides the convexity of the penalized loss in sparse regions to the greatest extent given certain thresholds for variable selection and unbiasedness. The PLUS computes multiple exact local minimizers of a possibly nonconvex penalized loss function in a certain main branch of the graph of critical points of the penalized loss. Its output is a continuous piecewise linear path encompassing from the origin for infinite penalty to a least squares solution for zero penalty. We prove that at a universal penalty level, the MC+ has high probability of matching the signs of the unknowns, and thus correct selection, without assuming the strong irrepresentable condition required by the LASSO. This selection consistency applies to the case of p >> n, and is proved to hold for exactly the MC+ solution among possibly many local minimizers. We prove that the MC+ attains certain minimax convergence rates in probability for the estimation of regression coefficients in e, balls. We use the SURE method to derive degrees of freedom and C-p-type risk estimates for general penalized LSE, including the LASSO and MC+ estimators, and prove their unbiasedness. Based on the estimated degrees of freedom, we propose an estimator of the noise level for proper choice of the penalty level. For full rank designs and general sub-quadratic penalties, we provide necessary and sufficient conditions for the continuity of the penalized LSE. Simulation results overwhelmingly support our claim Of Superior variable selection properties and demonstrate the computational efficiency of the proposed method.