LINEAR-MODEL SELECTION BY CROSS-VALIDATION

LINEAR-MODEL SELECTION BY CROSS-VALIDATION
复制标题

DOI:
10.2307/2290328
复制
发表时间:
1993-06-01
影响因子:
3.7
通讯作者:
SHAO, J
SHAO, J
中科院分区:
数学1区
文献类型:
--
作者:
SHAO, J

文献摘要

被引文献

相似文献

我们考虑在一类线性模型中选择具有最佳预测能力的模型的问题。流行的留一交叉验证方法与许多其他模型选择方法(如赤池信息准则(AIC)、C(p)和自举法)渐近等效,在选择具有最佳预测能力的模型的概率不收敛于1的意义上是渐近不一致的,因为观测总数n—>无穷。我们证明了留一交叉验证的不一致性可以通过使用留n(v)的交叉验证来纠正,n(v)是为验证保留的观测数,满足n(v)/n—> 1为n—>∞。这是一个有点令人震惊的发现,因为n(v)/n -> 1与交叉验证中流行的留一公式完全相反。提供了使用leave-n(v) out交叉验证方法的一些实际方面的动机,理由和讨论,并给出了模拟研究的结果。
We consider the problem of selecting a model having the best predictive ability among a class of linear models. The popular leave-one-out cross-validation method, which is asymptotically equivalent to many other model selection methods such as the Akaike information criterion (AIC), the C(p), and the bootstrap, is asymptotically inconsistent in the sense that the probability of selecting the model with the best predictive ability does not converge to 1 as the total number of observations n --> infinity. We show that the inconsistency of the leave-one-out cross-validation can be rectified by using a leave-n(v)-out cross-validation with n(v), the number of observations reserved for validation, satisfying n(v)/n --> 1 as n --> infinity. This is a somewhat shocking discovery, because n(v)/n --> 1 is totally opposite to the popular leave-one-out recipe in cross-validation. Motivations, justifications, and discussions of some practical aspects of the use of the leave-n(v)-out cross-validation method are provided, and results from a simulation study are presented.