Boosting with the L2 loss:: Regression and classification

Boosting with the L2 loss:: Regression and classification
复制标题

DOI:
10.1198/016214503000125
复制
发表时间:
2003-06-01
影响因子:
3.7
通讯作者:
Yu, B
Yu, B
中科院分区:
数学1区
文献类型:
--
作者:
Bühlmann, P;Yu, B

文献摘要

被引文献

相似文献

本文研究了一种计算简单的 boosting 变体 L(2)Boost,它是由具有 L-2 损失函数的函数梯度下降算法构造而成。与其他 boosting 算法一样,L(2)Boost 以迭代方式多次使用预先选择的拟合方法,称为学习器。基于L(2)Boost残差重新拟合的显式表达,在回归和分类方面详细研究了(对称)线性学习器的情况。特别是,当提升迭代 m 作为平滑或正则化参数时,会发现一种新的指数偏差-方差权衡,其中方差(复杂性)项随着 m 趋于无穷大而增长非常缓慢。当学习器是平滑样条时,最佳收敛率结果适用于回归和分类,并且增强的平滑样条甚至可以适应高阶、未知的平滑度。此外,推导了(平滑的)0-1 损失函数的简单展开,以揭示决策边界、偏差减少以及分类中加性偏差-方差分解的不可能性的重要性。最后,获得模拟和真实数据集结果来证明L(2)Boost的吸引力。特别是,我们证明了使用新颖的分量三次平滑样条的 L(2)Boosting 在存在高维预测变量的情况下既实用又有效。
This article investigates a computationally simple variant of boosting, L(2)Boost, which is constructed from a functional gradient descent algorithm with the L-2-loss function. Like other boosting algorithms, L(2)Boost uses many times in an iterative fashion a prechosen fitting method, called the learner. Based on the explicit expression of refitting of residuals of L(2)Boost, the case with (symmetric) linear learners is studied in detail in both regression and classification. In particular, with the boosting iteration m working as the smoothing or regularization parameter, a new exponential bias-variance trade-off is found with the variance (complexity) term increasing very slowly as m tends to infinity. When the learner is a smoothing spline, an optimal rate of convergence result holds for both regression and classification and the boosted smoothing spline even adapts to higher-order, unknown smoothness. Moreover, a simple expansion of a (smoothed) 0-1 loss function is derived to reveal the importance of the decision boundary, bias reduction, and impossibility of an additive bias-variance decomposition in classification. Finally, simulation and real dataset results are obtained to demonstrate the attractiveness of L(2)Boost. In particular, we demonstrate that L(2)Boosting with a novel component-wise cubic smoothing spline is both practical and effective in the presence of high-dimensional predictors.