Additive Logistic Regression : a Statistical

Additive Logistic Regression : a Statistical
复制标题

DOI:
--
复制
发表时间:
1998
期刊:
--
影响因子:
--
通讯作者:
T. Hastie
T. Hastie
中科院分区:
其他
文献类型:
--
作者:
T. Hastie

文献摘要

被引文献

相似文献

Booking(Freund&Schaire 1996,Schaire&Singer 1998)是分类方法论中最重要的发展之一。许多分类算法的性能通常可以通过将它们顺序地应用于输入数据的加权版本,并对由此产生的分类器序列进行加权多数投票来显著提高。我们表明,这一看似神秘的现象可以用众所周知的统计原理来理解,即加性建模和最大似然法。对于两类问题,Booking可以看作是对Logistic尺度上的加性建模的近似,使用最大伯努利似然作为准则。我们发展了更直接的近似,并表明它们表现出与Booking几乎相同的结果。导出了基于多项式似然的直接多类推广,其性能在大多数情况下可与最近提出的其他多类推广相媲美,但在某些情况下远远优于其他多类推广。我们建议对Booking稍作修改,以减少计算量,通常可减少10到50倍的计算量。最后,我们应用这些见解来产生一种提升决策树的替代公式。这种基于最优截断树归纳的方法通常会带来更好的性能,并且可以提供对聚集决策规则的可解释描述。它的计算速度也要快得多,使其更适合大规模数据挖掘应用。
Boosting (Freund & Schapire 1996, Schapire & Singer 1998) is one of the most important recent developments in classiication methodology. The performance of many classiication algorithms can often be dramatically improved by sequentially applying them to reweighted versions of the input data, and taking a weighted majority vote of the sequence of classiiers thereby produced. We show that this seemingly mysterious phenomenon can be understood in terms of well known statistical principles, namely additive modeling and maximum likelihood. For the two-class problem, boosting can be viewed as an approximation to additive modeling on the logistic scale using maximum Bernoulli likelihood as a criterion. We develop more direct approximations and show that they exhibit nearly identical results to boosting. Direct multi-class generalizations based on multinomial likelihood are derived that exhibit performance comparable to other recently proposed multi-class generalizations of boosting in most situations, and far superior in some. We suggest a minor modiication to boosting that can reduce computation, often by factors of 10 to 50. Finally, we apply these insights to produce an alternative formulation of boosting decision trees. This approach, based on best-rst truncated tree induction , often leads to better performance, and can provide interpretable descriptions of the aggregate decision rule. It is also much faster com-putationally making it more suitable to large scale data mining applications .