Improved boosting algorithms using confidence-rated predictions

Improved boosting algorithms using confidence-rated predictions
复制标题

DOI:
10.1023/a:1007614523901
复制
发表时间:
1999-12-01
期刊:
影响因子:
7.5
通讯作者:
Singer, Y
Singer, Y
中科院分区:
计算机科学3区
文献类型:
--
作者:
Schapire, RE;Singer, Y

文献摘要

被引文献

相似文献

我们描述了几个改进弗罗因德和Schapire的AdaBoost增强算法,特别是在设置中的假设可能会分配信心,他们的每个预测。我们给出了一个简化的分析AdaBoost在这种情况下,我们展示了如何使用这种分析来找到改进的参数设置,以及一个完善的标准训练弱假设。我们给出了一个具体的方法分配置信度的决策树的预测,一种方法密切相关的昆兰使用。这种方法还提出了一种用于生长决策树的技术,该技术与Kearns和Mansour提出的技术相同。接下来,我们将重点关注如何将新的提升算法应用于多类分类问题,特别是多标签的情况下,每个例子可能属于一个以上的类。针对这个问题,我们给出了两种增强方法,以及基于输出编码的第三种方法。其中之一导致了一种新的方法来处理单标签的情况下,这是简单的,但有效的技术所建议的Freund和Schapire。最后,我们给出了一些实验结果比较本文讨论的几个算法。
We describe several improvements to Freund and Schapire's AdaBoost boosting algorithm, particularly in a setting in which hypotheses may assign confidences to each of their predictions. We give a simplified analysis of AdaBoost in this setting, and we show how this analysis can be used to find improved parameter settings as well as a refined criterion for training weak hypotheses. We give a specific method for assigning confidences to the predictions of decision trees, a method closely related to one used by Quinlan. This method also suggests a technique for growing decision trees which turns out to be identical to one proposed by Kearns and Mansour. We focus next on how to apply the new boosting algorithms to multiclass classification problems, particularly to the multi-label case in which each example may belong to more than one class. We give two boosting methods for this problem, plus a third method based on output coding. One of these leads to a new method for handling the single-label case which is simpler but as effective as techniques suggested by Freund and Schapire. Finally, we give some experimental results comparing a few of the algorithms discussed in this paper.