Least angle regression

Least angle regression
复制标题

DOI:
10.1214/009053604000000067
复制
发表时间:
2004-04-01
影响因子:
4.5
通讯作者:
Tibshirani, R
Tibshirani, R
中科院分区:
数学1区
文献类型:
--
作者:
Efron, B;Hastie, T;Tibshirani, R

文献摘要

被引文献

相似文献

模型选择算法(如所有子集、向前选择和向后消除)的目的是基于将应用模型的同一组数据选择线性模型。通常,我们有大量可能的协变量,我们希望从中选择一个有效预测响应变量的简约集。最小角度回归(LARS)是一种新的模型选择算法,是传统前向选择方法的一种有用且不太贪婪的版本。三个主要属性推导:(1)一个简单的修改LARS算法实现了套索,一个有吸引力的版本的普通最小二乘约束的绝对回归系数的总和; LARS修改计算所有可能的套索估计为一个给定的问题,使用一个数量级的计算机时间比以前的方法。(2)一个不同的LARS修改有效地实现了前向逐步线性回归,另一个有前途的新的模型选择方法;这种连接解释了以前观察到的Lasso和Stagewise的类似数值结果,并帮助我们理解这两种方法的属性,这两种方法被视为更简单的LARS算法的约束版本。(3)LARS估计的自由度的一个简单的近似是可用的,从中我们得到一个C-P估计的预测误差,这允许一个原则性的选择可能的LARS估计的范围。LARS及其变体在计算上是高效的:本文描述了一种公开可用的算法,该算法只需要与应用于整组协变量的普通最小二乘法相同的计算量。
The purpose of model selection algorithms such as All Subsets, Forward Selection and Backward Elimination is to choose a linear model on the basis of the same set of data to which the model will be applied. Typically we have available a large collection of possible covariates from which we hope to select a parsimonious set for the efficient prediction of a response variable. Least Angle Regression (LARS), a new model selection algorithm, is a useful and less greedy version of traditional forward selection methods. Three main properties are derived: (1) A simple modification of the LARS algorithm implements the Lasso, an attractive version of ordinary least squares that constrains the sum of the absolute regression coefficients; the LARS modification calculates all possible Lasso estimates for a given problem, using an order of magnitude less computer time than previous methods. (2) A different LARS modification efficiently implements Forward Stagewise linear regression, another promising new model selection method; this connection explains the similar numerical results previously observed for the Lasso and Stagewise, and helps us understand the properties of both methods, which are seen as constrained versions of the simpler LARS algorithm. (3) A simple approximation for the degrees of freedom of a LARS estimate is available, from which we derive a C-p estimate of prediction error; this allows a principled choice among the range of possible LARS estimates. LARS and its variants are computationally efficient: the paper describes a publicly available algorithm that requires only the same order of magnitude of computational effort as ordinary least squares applied to the full set of covariates.