Training highly multiclass classifiers

Training highly multiclass classifiers
复制标题

DOI:
10.5555/2627435.2638582
复制
发表时间:
2014
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
M. Gupta;Samy Bengio;J. Weston
M. Gupta;Samy Bengio;J. Weston
中科院分区:
其他
文献类型:
--
作者:
M. Gupta;Samy Bengio;J. Weston

文献摘要

被引文献

相似文献

具有数千或更多类的分类问题通常具有很大范围的类混淆性,并且我们表明,在训练期间最小化的经验损失中,更容易混淆的类增加了更多的噪声。我们提出了一种在线解决方案,减少了高度混淆的类对训练分类器参数的影响,并将训练重点放在训练中任何给定时间都更容易区分的类对上。我们还表明,最近提出的用于自动减少凸随机梯度下降步长的adagrad方法也可以有效地应用于监督降维和线性分类器的非凸联合训练,就像Wsabie所做的那样。在ImageNet基准数据集和拥有15,000到97,000个类的专有图像识别问题上进行的实验表明,与一对一的线性支持向量机和Wsabie相比,分类精度有了很大提高。
Classification problems with thousands or more classes often have a large range of class-confusabilities, and we show that the more-confusable classes add more noise to the empirical loss that is minimized during training. We propose an online solution that reduces the effect of highly confusable classes in training the classifier parameters, and focuses the training on pairs of classes that are easier to differentiate at any given time in the training. We also show that the adagrad method, recently proposed for automatically decreasing step sizes for convex stochastic gradient descent, can also be profitably applied to the nonconvex joint training of supervised dimensionality reduction and linear classifiers as done in Wsabie. Experiments on ImageNet benchmark data sets and proprietary image recognition problems with 15,000 to 97,000 classes show substantial gains in classification accuracy compared to one-vs-all linear SVMs and Wsabie.