Variable selection - A review and recommendations for the practicing statistician.

Variable selection - A review and recommendations for the practicing statistician.
复制标题

DOI:
10.1002/bimj.201700067
复制
发表时间:
2018-05
期刊:
Biometrical journal. Biometrische Zeitschrift
影响因子:
--
通讯作者:
Dunkler D
Dunkler D
中科院分区:
其他
文献类型:
--
作者:
Heinze G;Wallisch C;Dunkler D

文献摘要

参考文献

被引文献

相似文献

统计模型支持医学研究,通过促进独立变量条件下的个体化结果预测,或通过估计协变量调整后的风险因素的影响。如果要考虑的自变量集是固定的且很小,则统计模型理论是完善的。因此,我们可以假设效应估计是无偏的,并且通常的置信区间估计方法是有效的。然而,在日常工作中,并不事先知道哪些协变量应该包括在模型中,我们经常遇到候选变量的数量在10-30之间。这个数字通常太大,无法在统计模型中考虑。我们提供了各种可用的变量选择方法的概述,这些方法基于显著性或信息标准、惩罚似然、估计值变化标准、背景知识或其组合。这些方法通常是在线性回归模型的背景下开发的,然后转移到更广义的线性模型或删失生存数据模型。变量选择,特别是如果用于解释性建模,其中效应估计值是主要关注的,可能会损害最终模型的稳定性,回归系数的无偏性以及p值或置信区间的有效性。因此,我们为实践统计学家在一般(低维)建模问题中应用变量选择方法以及进行稳定性调查和推理提供了实用的建议。我们还提出了一些数量的基础上restaurant的整个变量选择过程中定期报告的软件包提供自动变量选择算法。
Statistical models support medical research by facilitating individualized outcome prognostication conditional on independent variables or by estimating effects of risk factors adjusted for covariates. Theory of statistical models is well‐established if the set of independent variables to consider is fixed and small. Hence, we can assume that effect estimates are unbiased and the usual methods for confidence interval estimation are valid. In routine work, however, it is not known a priori which covariates should be included in a model, and often we are confronted with the number of candidate variables in the range 10–30. This number is often too large to be considered in a statistical model. We provide an overview of various available variable selection methods that are based on significance or information criteria, penalized likelihood, the change‐in‐estimate criterion, background knowledge, or combinations thereof. These methods were usually developed in the context of a linear regression model and then transferred to more generalized linear models or models for censored survival data. Variable selection, in particular if used in explanatory modeling where effect estimates are of central interest, can compromise stability of a final model, unbiasedness of regression coefficients, and validity of p‐values or confidence intervals. Therefore, we give pragmatic recommendations for the practicing statistician on application of variable selection methods in general (low‐dimensional) modeling problems and on performing stability investigations and inference. We also propose some quantities based on resampling the entire variable selection process to be routinely reported by software packages offering automated variable selection algorithms.
DOI: 10.1186/1751-0473-3-17
发表时间: 2008-12-16
影响因子: --
作者:
Bursac, Zoran;Gauss, C Heath;Williams, David Keith;Hosmer, David W
通讯作者: Hosmer, David W
DOI: 10.1093/ije/29.1.158
发表时间: 2000-02-01
影响因子: 7.7
作者:
Greenland, S
通讯作者: Greenland, S
DOI: 10.1093/biomet/asr041
发表时间: 2011-12-01
期刊: BIOMETRIKA
影响因子: 2.7
作者:
De Luna, Xavier;Waernbaum, Ingeborg;Richardson, Thomas S.
通讯作者: Richardson, Thomas S.
DOI: 10.1214/ss/1009213726
发表时间: 2001-08-01
影响因子: 5.7
作者:
Breiman, L
通讯作者: Breiman, L
DOI: 10.1198/016214503000125
发表时间: 2003-06-01
影响因子: 3.7
作者:
Bühlmann, P;Yu, B
通讯作者: Yu, B