Variable selection in linear regression

Variable selection in linear regression
复制标题

DOI:
10.1177/1536867x1101000407
复制
发表时间:
2010-01-01
期刊:
影响因子:
4.8
通讯作者:
Sheather, Simon
Sheather, Simon
中科院分区:
数学3区
文献类型:
--
作者:
Lindsey, Charles;Sheather, Simon

文献摘要

被引文献

相似文献

我们提出了一个新的 Stata 程序 vselect,它可以帮助用户在执行线性回归后执行变量选择。提供了逐步方法的选项,例如前向选择和后向消除。用户可以指定Mallows的C-p、Akaike的信息准则、Akaike的校正信息准则、贝叶斯信息准则或调整的R-2作为选择的信息准则。当用户指定最佳子集选项时,跳跃式算法(Furnival 和 Wilson,Technometrics 16:499-511)用于确定每个预测变量大小的最佳子集。为每个子集报告所有前面提到的信息标准。我们还提供仅对某些预测变量进行变量选择的选项(如 [R] Nestreg 中)并支持加权线性回归。所有选项均在具有不同数量预测变量的真实数据集上进行演示。
We present a new Stata program, vselect, that helps users perform variable selection after performing a linear regression. Options for stepwise methods such as forward selection and backward elimination are provided. The user may specify Mallows's C-p, Akaike's information criterion, Akaike's corrected information criterion, Bayesian information criterion, or R-2 adjusted as the information criterion for the selection. When the user specifies the best subset option, the leaps-and-bounds algorithm (Furnival and Wilson, Technometrics 16: 499-511) is used to determine the best subsets of each predictor size. All the previously mentioned information criteria are reported for each of these subsets. We also provide options for doing variable selection only on certain predictors (as in [R] nestreg) and support for weighted linear regression. All options are demonstrated on real datasets with varying numbers of predictors.