Greedy function approximation: A gradient boosting machine

Greedy function approximation: A gradient boosting machine
复制标题

DOI:
10.1214/aos/1013203451
复制
发表时间:
2001-10-01
影响因子:
4.5
通讯作者:
Friedman, JH
Friedman, JH
中科院分区:
数学1区
文献类型:
--
作者:
Friedman, JH

文献摘要

被引文献

相似文献

函数估计/逼近是从函数空间而不是参数空间的数值优化的角度来考虑的。在阶段性加性扩展和最陡下降最小化之间建立了联系。基于任意拟合准则,发展了一种适用于加性展开的广义梯度下降“助推”范式。给出了用于回归的最小二乘、最小绝对偏差和Huber-M损失函数以及用于分类的多类Logistic似然估计的具体算法。针对单个加性分量是回归树的特定情况,导出了特殊的增强,并给出了用于解释这种“TreeBoost”模型的工具。回归树的梯度增强为回归和分类产生了具有竞争力的、高度健壮的、可解释的过程,尤其适合于破坏不太干净的数据。文中还讨论了这种方法与Freund和Shapire以及Friedman、Hastie和Tibshiani的增压方法之间的联系。
Function estimation/approximation is viewed from the perspective of numerical optimization in function space, rather than parameter space. A connection is made between stagewise additive expansions and steepest-descent minimization. A general gradient descent "boosting" paradigm is developed for additive expansions based on any fitting criterion. Specific algorithms are presented for least-squares, least absolute deviation, and Huber-M loss functions for regression, and multiclass logistic likelihood for classification. Special enhancements are derived for the particular case where the individual additive components are regression trees, and tools for interpreting such "TreeBoost" models are presented. Gradient boosting of regression trees produces competitive, highly robust, interpretable procedures for both regression and classification, especially appropriate for ruining less than clean data. Connections between this approach and the boosting methods of Freund and Shapire and Friedman, Hastie and Tibshirani are discussed.