Convexity, classification, and risk bounds

Convexity, classification, and risk bounds
复制标题

DOI:
10.1198/016214505000000907
复制
发表时间:
2006-03-01
影响因子:
3.7
通讯作者:
McAuliffe, JD
McAuliffe, JD
中科院分区:
数学1区
文献类型:
--
作者:
Bartlett, PL;Jordan, MI;McAuliffe, JD

文献摘要

被引文献

相似文献

在机器学习文献中开发的许多分类算法,包括支持向量机和boosting,可以被视为最小对比度方法,其最小化0-1损失函数的凸代理。凸性使得这些算法计算效率高。然而,使用替代,统计的后果,必须对凸性的计算优点进行平衡。为了研究这些问题,我们提供了一个一般的定量关系之间的风险评估使用0-1损失和风险评估使用任何非负的替代损失函数。我们发现,这种关系给出了非平凡的上限下的最弱的可能条件下的损失函数的过度风险,它满足分类的Fisher一致性的逐点形式。该关系是基于一个简单的变分变换的损失函数,很容易计算在许多应用中。我们还提出了一个改进的版本,在低噪声的情况下,这一结果。并表明,在这种情况下,严格凸损失函数导致更快的收敛速度的风险比标准的一致收敛参数所暗示的。最后,我们提出的应用程序,我们的研究结果估计的收敛速度的函数类,缩放凸包的有限维基类,与各种常用的损失函数。
Many of the classification algorithms developed in the machine learning literature, including the support vector machine and boosting, can be viewed as minimum contrast methods that minimize a convex surrogate of the 0-1 loss function. The convexity makes these algorithms computationally efficient. The use of a surrogate, however, has statistical consequences that must be balanced against the computational virtues of convexity. To study these issues, we provide a general quantitative relationship between the risk as assessed using the 0-1 loss and the risk as assessed using any nonnegative surrogate loss function. We show that this relationship gives nontrivial upper bounds on excess risk under the weakest possible condition on the loss function-that it satisfies a pointwise form of Fisher consistency for classification. The relationship is based on a simple variational transformation of the loss function that is easy to compute in many applications. We also present a refined version of this result in the case of low noise. and show that in this case, strictly convex loss functions lead to faster rates of convergence of the risk than would be implied by standard uniform convergence arguments. Finally, we present applications of our results to the estimation of convergence rates in function classes that are scaled convex hulls of a finite-dimensional base class, with a variety of commonly used loss functions.