Training highly multiclass classifiers
Training highly multiclass classifiers
复制标题
DOI:
10.5555/2627435.2638582
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
M. Gupta;Samy Bengio;J. Weston
中科院分区:
文献类型:
--
作者:
M. Gupta;Samy Bengio;J. Weston
Classification problems with thousands or more classes often have a large range of class-confusabilities, and we show that the more-confusable classes add more noise to the empirical loss that is minimized during training. We propose an online solution that reduces the effect of highly confusable classes in training the classifier parameters, and focuses the training on pairs of classes that are easier to differentiate at any given time in the training. We also show that the adagrad method, recently proposed for automatically decreasing step sizes for convex stochastic gradient descent, can also be profitably applied to the nonconvex joint training of supervised dimensionality reduction and linear classifiers as done in Wsabie. Experiments on ImageNet benchmark data sets and proprietary image recognition problems with 15,000 to 97,000 classes show substantial gains in classification accuracy compared to one-vs-all linear SVMs and Wsabie.