An efficient method to estimate bagging's generalization error

An efficient method to estimate bagging's generalization error
复制标题

DOI:
10.1023/a:1007519102914
复制
发表时间:
1999-04-01
期刊:
影响因子:
7.5
通讯作者:
Macready, WG
Macready, WG
中科院分区:
计算机科学3区
文献类型:
--
作者:
Wolpert, DH;Macready, WG

文献摘要

被引文献

相似文献

Bagging(Breiman,1994 a)是一种尝试通过使用训练集的自举复制来提高学习算法性能的技术(Efron & Tibshirani,1993,Efron,1979)。通过交叉验证来估计测试集上的结果泛化误差的计算要求通常是禁止的,对于留一交叉验证,需要在我的时间的顺序上训练底层算法,其中m是训练集的大小,v是重复的数量。本文提出了几种技术,用于估计装袋学习算法的泛化误差,而无需调用更多的底层学习算法的训练(超出装袋本身),这是基于交叉验证的估计所需的。这些技术都利用了偏差-方差分解(Geman,Bienenstock & Doursat,1992,Wolpert,1996)。我们最好的估计器也利用了堆叠(Wolpert,1992)。在这里报告的一组实验中,它被发现比替代的基于交叉验证的袋装算法的误差估计和基于交叉验证的底层算法的误差估计更准确。这种改进对于小型测试集尤其明显。这提出了一个新的理由使用装袋-更准确的估计泛化误差比没有装袋。
Bagging (Breiman, 1994a) is a technique that tries to improve a learning algorithm's performance by using bootstrap replicates of the training set (Efron & Tibshirani, 1993, Efron, 1979). The computational requirements for estimating the resultant generalization error on a test set by means of cross-validation are often prohibitive, for leave-one-out cross-validation one needs to train the underlying algorithm on the order of my times, where m is the size of the training set and v is the number of replicates. This paper presents several techniques for estimating the generalization error of a bagged learning algorithm without invoking yet more training of the underlying learning algorithm (beyond that of the bagging itself), as is required by cross-validation-based estimation. These techniques all exploit the bias-variance decomposition (Geman, Bienenstock & Doursat, 1992, Wolpert, 1996). The best of our estimators also exploits stacking (Wolpert, 1992). In a set of experiments reported here, it was found to be more accurate than both the alternative cross-validation-based estimator of the bagged algorithm's error and the cross-validation-based estimator of the underlying algorithm's error. This improvement was particularly pronounced for small test sets. This suggests a novel justification for using bagging- more accurate estimation of the generalization error than is possible without bagging.