Incremental Learning of Random Forests for Large-Scale Image Classification

Incremental Learning of Random Forests for Large-Scale Image Classification
复制标题

DOI:
10.1109/tpami.2015.2459678
复制
发表时间:
2016-03
影响因子:
23.6
通讯作者:
M. Ristin;M. Guillaumin;Juergen Gall;L. Gool
M. Ristin;M. Guillaumin;Juergen Gall;L. Gool
中科院分区:
计算机科学1区
文献类型:
--
作者:
M. Ristin;M. Guillaumin;Juergen Gall;L. Gool

文献摘要

被引文献

相似文献

像ImageNet这样的大型图像数据集或像Flickr这样的开放式照片网站正在揭示图像分类的新挑战,这些挑战在较小的固定集中并不明显。特别是,动态增长的数据集的有效处理,不仅是训练数据的数量,而且类的数量也随着时间的推移而增加,是一个相对未被探索的问题。在这个具有挑战性的环境中,我们研究了随机森林(RF)的两种变体在四种策略下的表现,以合并新类,同时避免从头开始重新训练RF。各种策略在分类精度和计算效率之间做出了不同的权衡。在我们广泛的实验中,我们表明两个RF变体,一个基于最近类均值分类器,另一个基于支持向量机,都优于传统RF,并且非常适合增量学习新类。特别是,我们表明,与从完整数据中训练相比,最初仅用10个类训练的RFs可以扩展到1,000个类,其准确性损失是可以接受的,并且与为每批新类重新训练相比,节省了大量的计算量。
Large image datasets such as ImageNet or open-ended photo websites like Flickr are revealing new challenges to image classification that were not apparent in smaller, fixed sets. In particular, the efficient handling of dynamically growing datasets, where not only the amount of training data but also the number of classes increases over time, is a relatively unexplored problem. In this challenging setting, we study how two variants of Random Forests (RF) perform under four strategies to incorporate new classes while avoiding to retrain the RFs from scratch. The various strategies account for different trade-offs between classification accuracy and computational efficiency. In our extensive experiments, we show that both RF variants, one based on Nearest Class Mean classifiers and the other on SVMs, outperform conventional RFs and are well suited for incrementally learning new classes. In particular, we show that RFs initially trained with just 10 classes can be extended to 1,000 classes with an acceptable loss of accuracy compared to training from the full data and with great computational savings compared to retraining for each new batch of classes.