Variable selection via nonconcave penalized likelihood and its oracle properties

Variable selection via nonconcave penalized likelihood and its oracle properties
复制标题

DOI:
10.1198/016214501753382273
复制
发表时间:
2001-12-01
影响因子:
3.7
通讯作者:
Li, RZ
Li, RZ
中科院分区:
数学1区
文献类型:
--
作者:
Fan, JQ;Li, RZ

文献摘要

被引文献

相似文献

变量选择是高维统计建模的基础,包括非参数回归。目前使用的许多方法都是逐步选择程序,这种方法计算成本很高,并且忽略了变量选择过程中的随机误差。本文提出了惩罚似然方法来处理这类问题。所提出的方法同时选择变量和估计系数。因此,它们使我们能够为估计的参数构造可信区间。所提出的方法与其他方法的不同之处在于罚函数是对称的,在(0,无穷大)上是非凹的,并且在原点具有奇性以产生稀疏解。此外,罚函数应该以一个常数为界,以减少偏差,并满足某些条件以产生连续解。提出了一种优化惩罚似然函数的新算法。所提出的想法具有广泛的适用性。它们可以很容易地应用于各种参数模型,如广义线性模型和稳健回归模型。它们还可以通过使用小波和样条线轻松地应用于非参数建模。给出了惩罚似然估计的收敛速度。此外,在适当选择正则化参数的情况下,我们证明了所提出的估计器在变量选择上的表现与Oracle过程一样好;即,如果已知正确的子模型,它们的工作也一样好。仿真结果表明,与其他变量选择方法相比,新提出的方法具有更好的性能。更重要的是。经检验,标准误差公式具有足够的精度,可以满足实际应用的需要。
Variable selection is fundamental to high-dimensional statistical modeling, including nonparametric regression. Many approaches in use are stepwise selection procedures, which can be computationally expensive and ignore stochastic errors in the variable selection process. In this article, penalized likelihood approaches are proposed to handle these kinds of problems. The proposed methods select variables and estimate coefficients simultaneously. Hence they enable us to construct confidence intervals for estimated parameters. The proposed approaches are distinguished from others in that the penalty functions are symmetric, nonconcave on (0, infinity), and have singularities at the origin to produce sparse solutions. Furthermore, the penalty functions should be bounded by a constant to reduce bias and satisfy certain conditions to yield continuous solutions. A new algorithm is proposed for optimizing penalized likelihood functions. The proposed ideas are widely applicable. They are readily applied to a variety of parametric models such as generalized linear models and robust regression models. They can also be applied easily to nonparametric modeling by using wavelets and splines. Rates of convergence of the proposed penalized likelihood estimators are established. Furthermore, with proper choice of regularization parameters, we show that the proposed estimators perform as well as the oracle procedure in variable selection; namely, they work as well as if the correct submodel were known. Our simulation shows that the newly proposed methods compare favorably with other variable selection techniques. Furthermore. the standard error formulas are tested to be accurate enough for practical applications.