STACKED GENERALIZATION

STACKED GENERALIZATION
复制标题

DOI:
10.1016/s0893-6080(05)80023-1
复制
发表时间:
1992-01-01
期刊:
影响因子:
7.8
通讯作者:
WOLPERT, DH
WOLPERT, DH
中科院分区:
计算机科学1区
文献类型:
--
作者:
WOLPERT, DH

文献摘要

被引文献

相似文献

本文介绍了堆叠泛化,这是一种最小化一个或多个泛化器泛化错误率的方案。堆叠泛化的工作原理是推导泛化者(S)相对于所提供的学习集的偏差。这种推论通过在第二个空间中进行推广,该空间的输入是(例如)当用学习集的一部分教授原始泛化者的猜测时,并试图猜测它的其余部分,而其输出是(例如)正确的猜测。当与多个泛化器一起使用时,堆叠泛化可以被视为更复杂的交叉验证版本,它利用一种比交叉验证粗略的赢家通吃更复杂的策略来组合各个泛化器。当与单个泛化器一起使用时,堆叠泛化是一种方案,用于估计(然后纠正)已在特定学习集上训练并随后提出特定问题的泛化器的错误。在介绍了堆叠泛化并证明了它的使用之后,本文给出了两个数值实验。第一个演示了堆叠泛化如何在一组单独的泛化器的基础上改进,用于netTalk将文本转换为音素的任务。第二个演示了堆叠泛化如何提高单个曲面修配者的性能。结合文献中的其他实验证据,支持交叉验证的常见论点,以及本文提出的抽象理由,结论是对于几乎任何现实世界的泛化问题,都应该使用某种版本的堆叠泛化来最小化泛化错误率。本文最后讨论了堆叠泛化的一些变化,以及它如何涉及到其他领域,如混沌理论。
This paper introduces stacked generalization, a scheme for minimizing the generalization error rate of one or more generalizers. Stacked generalization works by deducing the biases of the generalizer(s) with respect to a provided learning set. This deduction proceeds by generalizing in a second space whose inputs are (for example) the guesses of the original generalizers when taught with part of the learning set and trying to guess the rest of it, and whose output is (for example) the correct guess. When used with multiple generalizers, stacked generalization can be seen as a more sophisticated version of cross-validation, exploiting a strategy more sophisticated than cross-validation's crude winner-takes-all for combining the individual generalizers. When used with a single generalizer, stacked generalization is a scheme for estimating (and then correcting for) the error of a generalizer which has been trained on a particular learning set and then asked a particular question. After introducing stacked generalization and justifying its use, this paper presents two numerical experiments. The first demonstrates how stacked generalization improves upon a set of separate generalizers for the NETtalk task of translating text to phonemes. The second demonstrates how stacked generalization improves the performance of a single surface-fitter. With the other experimental evidence in the literature, the usual arguments supporting cross-validation, and the abstract justifications presented in this paper, the conclusion is that for almost any real-world generalization problem one should use some version of stacked generalization to minimize the generalization error rate. This paper ends by discussing some of the variations of stacked generalization, and how it touches on other fields like chaos theory.