Learn from each other to Classify better: Cross-layer mutual attention learning for fine-grained visual classification

Learn from each other to Classify better: Cross-layer mutual attention learning for fine-grained visual classification
复制标题

DOI:
10.1016/j.patcog.2023.109550
复制
发表时间:
2023-03
期刊:
Pattern Recognit.
影响因子:
--
通讯作者:
Dichao Liu;Long Zhao;Yu Wang;Jien Kato
Dichao Liu;Long Zhao;Yu Wang;Jien Kato
中科院分区:
其他
文献类型:
--
作者:
Dichao Liu;Long Zhao;Yu Wang;Jien Kato

文献摘要

相似文献

细粒度视觉分类(FGVC)既有价值,又具有挑战性。FGVC的困难主要在于其固有的类间相似性、类内变异和训练数据有限。此外,随着深卷积神经网络的普及,研究人员主要使用深层的、抽象的、语义的信息来表示模糊逻辑,而忽略了浅层的、详细的信息。本文提出了一种跨层交互注意学习网络(CMAL-Net)来解决上述问题。具体地说,这项工作将CNN的浅层到深层视为通晓不同视角的“专家”。我们让每个专家给出一个类别预测和一个指示所发现线索的关注区域。关注区域被视为专家之间的信息载体,带来了三个好处:(I)帮助模型专注于区分区域;(I I)提供更多的训练数据;(I I I)允许专家相互学习,以提高整体性能。CMAL-Net在三个竞争数据集上实现了最先进的性能:FGVC-Airline、Stanford Cars和Food-11。源代码可在https://github.上找到Com/刘迪超/CMAL
Fine-grained visual classification (FGVC) is valuable yet challenging. The difficulty of FGVC mainly lies in its intrinsic inter-class similarity, intra-class variation, and limited training data. Moreover, with the popularity of deep convolutional neural networks, researchers have mainly used deep, abstract, semantic information for FGVC, while shallow, detailed information has been neglected. This work proposes a cross-layer mutual attention learning network (CMAL-Net) to solve the above problems. Specifically, this work views the shallow to deep layers of CNNs as “experts” knowledgeable about different perspectives. We let each expert give a category prediction and an attention region indicating the found clues. Attention regions are treated as information carriers among experts, bringing three benefits:(i) helping the model focus on discriminative regions;(i i) providing more training data;(i i i) allowing experts to learn from each other to improve the overall performance. CMAL-Net achieves state-of-the-art performance on three competitive datasets: FGVC-Aircraft, Stanford Cars, and Food-11. The source code is available at https://github. com/Dichao-Liu/CMAL