Cascaded one-vs-rest detection network for fine-grained recognition without part annotations

Cascaded one-vs-rest detection network for fine-grained recognition without part annotations
复制标题

DOI:
10.1007/s11042-018-5875-y
复制
发表时间:
2017-02
影响因子:
3.6
通讯作者:
Long Chen;Shengke Wang;K. Lam;Huiyu Zhou;Muwei Jian;Junyu Dong
Long Chen;Shengke Wang;K. Lam;Huiyu Zhou;Muwei Jian;Junyu Dong
中科院分区:
计算机科学4区
文献类型:
--
作者:
Long Chen;Shengke Wang;K. Lam;Huiyu Zhou;Muwei Jian;Junyu Dong

文献摘要

相似文献

由于类别内差异较小,细粒度识别是一项具有挑战性的任务。大多数性能最佳的细粒度识别方法都利用对象的一部分来获得更好的性能。因此,需要计算量极大的部件注释。在本文中,我们提出了一种用于细粒度识别的新型级联深度 CNN 检测框架,该框架经过训练可以检测整个对象而不考虑各个部分。尽管如此,目前大多数表现最好的检测网络都使用 N + 1 类(N 个对象类别加上背景)softmax 损失。训练样本较多的背景类别主导了特征学习进度,而这些特征不适合样本较少的对象分类。为了解决这个问题,我们在这里介绍两种策略:1)我们利用级联结构来消除背景。 2)我们引入了一种新颖的一对一损失函数来捕获来自不同从属类别的更多微小差异。实验表明,我们提出的识别框架在 CUB-200-2011 Bird 数据集上实现了与最先进的、无部分、细粒度的识别方法相当的性能。同时,我们的方法优于大多数现有的基于零件注释的方法,并且在训练阶段不需要零件注释,同时在测试阶段不需要任何注释。
Fine-grained recognition is a challenging task due to small intra-category variances. Most of the top-performing fine-grained recognition methods leverage parts of objects for better performance. Therefore, part annotations which are extremely computationally expensive are required. In this paper, we propose a novel cascaded deep CNN detection framework for fine-grained recognition which is trained to detect a whole object without considering parts. Nevertheless, most of the current top-performing detection networks use N + 1 class (N object categories plus background) softmax loss. The background category with much more training samples dominates the feature learning progress where the features are not suitable for object categorisation with fewer samples. To address this issue, we here introduce two strategies: 1) We leverage a cascaded structure to eliminate the background. 2) We introduce a novel one-vs-rest loss function to capture more minute variances from different subordinate categories. Experiments show that our proposed recognition framework achieves comparable performance against the state-of-the-art, part-free, fine-grained recognition methods on the CUB-200-2011 Bird dataset. Meanwhile, our method outperforms most of the existing part annotation based methods and does not need part annotations at the training stage whilst being free from any annotations at the test stage.