Estimating classification error rate: Repeated cross-validation, repeated hold-out and bootstrap

Estimating classification error rate: Repeated cross-validation, repeated hold-out and bootstrap
复制标题

DOI:
10.1016/j.csda.2009.04.009
复制
发表时间:
2009-09-01
影响因子:
1.8
通讯作者:
Kim, Ji-Hyun
Kim, Ji-Hyun
中科院分区:
数学3区
文献类型:
--
作者:
Kim, Ji-Hyun

文献摘要

被引文献

相似文献

我们考虑了在给定训练样本上构造的分类器的精度估计。众所周知,天真的再替代估计存在向下偏差的问题。解决这一偏见问题的传统方法是交叉验证。Bootstrap是降低交叉验证高度可变性的另一种方法。但是,直接比较交叉验证和自举两种估计量是不公平的,因为后一种估计量需要更大的计算量。我们进行了一项实证研究,以比较.632+Bootstrap估计与重复10倍交叉验证和重复三分之一坚持估计。所有的估计器都被设置为需要大约相同的计算量。在模拟研究中,当分类器高度适应训练样本时,重复10次交叉验证估计器被发现具有比.632+Bootstrap估计器更好的性能。我们还发现.632+Bootstrap估计器对于大样本和小样本都存在偏差问题。(C)2009爱思唯尔B.V.保留所有权利。
We consider the accuracy estimation of a classifier constructed on a given training sample. The naive resubstitution estimate is known to have a downward bias problem. The traditional approach to tackling this bias problem is cross-validation. The bootstrap is another way to bring down the high variability of cross-validation. But a direct comparison of the two estimators, cross-validation and bootstrap, is not fair because the latter estimator requires much heavier computation. We performed an empirical study to compare the .632+ bootstrap estimator with the repeated 10-fold cross-validation and the repeated one-third holdout estimator. All the estimators were set to require about the same amount of computation. In the simulation study, the repeated 10-fold cross-validation estimator was found to have better performance than the .632+ bootstrap estimator when the classifier is highly adaptive to the training sample. We have also found that the .632+ bootstrap estimator suffers from a bias problem for large samples as well as for small samples. (C) 2009 Elsevier B.V. All rights reserved.