Degrees of freedom and model search

Degrees of freedom and model search
复制标题

自由度和模型搜索

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
R. Tibshirani
R. Tibshirani
中科院分区:
--
文献类型:
--
作者:
R. Tibshirani

文献摘要

被引文献

相似文献

自由度是统计建模中的一个基本概念,因为它提供了对给定过程执行的拟合量的定量描述。但是,尽管它在统计学中发挥着重要作用,但它的行为并没有得到完全的理解,即使在一些基本的环境中也是如此。例如,它可能看起来直观明显,子集大小为k的最佳子集选择拟合具有大于k的自由度,但这尚未得到正式验证,也没有被精确研究。在大,目前的文件是出于这个问题,我们推导出一个精确的表达式的自由度的最佳子集选择在限制设置(正交预测变量)。沿着的方式,我们开发了一个概念,我们命名为“搜索自由度”;直观地说,对于执行变量选择的自适应回归过程,这是我们完全归因于模型选择机制的(总)自由度的一部分。最后,我们建立了一个适度的扩展Stein的公式,涵盖不连续的功能,并讨论其潜在的作用,在自由度和搜索自由度的计算。
Degrees of freedom is a fundamental concept in statistical modeling, as it provides a quantitative description of the amount of fitting performed by a given procedure. But, despite this fundamental role in statistics, its behavior is not completely well-understood, even in somewhat basic settings. For example, it may seem intuitively obvious that the best subset selection fit with subset size k has degrees of freedom larger than k, but this has not been formally verified, nor has is been precisely studied. At large, the current paper is motivated by this problem, and we derive an exact expression for the degrees of freedom of best subset selection in a restricted setting (orthogonal predictor variables). Along the way, we develop a concept that we name “search degrees of freedom”; intuitively, for adaptive regression procedures that perform variable selection, this is a part of the (total) degrees of freedom that we attribute entirely to the model selection mechanism. Finally, we establish a modest extension of Stein’s formula to cover discontinuous functions, and discuss its potential role in degrees of freedom and search degrees of freedom calculations.