Additive logistic regression: A statistical view of boosting

Additive logistic regression: A statistical view of boosting
复制标题

DOI:
10.1214/aos/1016218223
复制
发表时间:
2000-04-01
影响因子:
4.5
通讯作者:
Tibshirani, R
Tibshirani, R
中科院分区:
数学1区
文献类型:
--
作者:
Friedman, J;Hastie, T;Tibshirani, R

文献摘要

被引文献

相似文献

Boosting是分类方法学最重要的近期发展之一。Boosting通过顺序地将分类算法应用于训练数据的重新加权版本,然后对由此产生的分类器序列进行加权多数投票来工作。对于许多分类算法来说,这种简单的策略可以显著提高性能。我们表明,这个看似神秘的现象可以理解的众所周知的统计原则,即添加剂建模和最大似然。对于两类问题,提升可以被视为使用最大伯努利似然作为标准的逻辑尺度上的加性建模的近似。我们开发了更直接的近似,并表明他们表现出几乎相同的结果,以提高。直接多类概括的基础上多项式的可能性,表现出性能相媲美的其他最近提出的多类概括的提高在大多数情况下,远远优于上级在一些。我们建议对boosting进行一个小的修改,可以减少计算量,通常是10到50倍。最后,我们应用这些见解来生成一种替代的提升决策树的公式。这种方法,基于最佳优先截断树归纳,往往会导致更好的性能,并可以提供可解释的描述的聚合决策规则。它的计算速度也更快,使其更适合大规模的数据挖掘应用。
Boosting is one of the most important recent developments in classification methodology. Boosting works by sequentially applying a classification algorithm to reweighted Versions of the training data and then taking a weighted majority vote of the sequence of classifiers thus produced. For many classification algorithms, this simple strategy results in dramatic improvements in performance. We show that this seemingly mysterious phenomenon can be understood in terms of well-known statistical principles, namely additive modeling and maximum likelihood. For the two-class problem, boosting can be viewed as an approximation to additive modeling on the logistic scale using maximum Bernoulli likelihood as a criterion. We develop more direct approximations and show that they exhibit nearly identical results to boosting. Direct multiclass generalizations based on multinomial likelihood are derived that exhibit performance comparable to other recently proposed multiclass generalizations of boosting in most situations, and far superior in some. We suggest a minor modification to boosting that can reduce computation, often by factors of 10 to 50. Finally, we apply these insights to produce an alternative formulation of boosting decision trees. This approach, based on best-first truncated tree induction, often leads to better performance, and can provide interpretable descriptions of the aggregate decision rule. It is also much faster computationally, making it more suitable to large-scale data mining applications.