Bi-Modal Progressive Mask Attention for Fine-Grained Recognition

Bi-Modal Progressive Mask Attention for Fine-Grained Recognition
复制标题

用于细粒度识别的双模态渐进掩模注意

DOI:
10.1109/tip.2020.2996736
复制
发表时间:
2020-01-01
影响因子:
10.6
通讯作者:
Lu, Jianfeng
Lu, Jianfeng
中科院分区:
计算机科学1区
文献类型:
--
作者:
Song, Kaitao;Wei, Xiu-Shen;Lu, Jianfeng

文献摘要

被引文献

相似文献

传统的细粒度图像识别需要根据原始图像下的视觉线索来区分不同的从属类别(如鸟类)。由于既有小的类间差异,也有大的类内差异,需要捕捉这些子类别之间的细微差异,这对于细粒度识别来说是至关重要的,但也是具有挑战性的。近年来,语言情态聚合被证明是一种成功的提高视觉识别能力的技术。本文介绍了一种端到端可训练的渐进掩模注意(PMA)模型,该模型利用视觉和语言两种方式进行细粒度识别。我们的双通道PMA模型不仅可以通过基于掩模的方式逐级捕捉视觉通道中最具区分性的部分,而且可以在交互对齐范式下从语言通道中探索视觉外的知识。具体地说,在每个阶段,一个自我注意模块被提出来关注图像或文本描述中的关键补丁。此外,还设计了查询关系模块来捕捉文本中的关键词/短语,从而在两个通道之间架起一座桥梁。随后,从多个阶段学习的双通道表示被聚合为最终的特征用于识别。我们的双模PMA模型只需要原始图像和原始文本描述,而不需要图像中的边界框/部件注释或文本中的关键词注释。通过在细粒度基准数据集上的综合实验,我们证明了该方法无论是在视觉和语言双通道还是在单一视觉通道上都取得了优于竞争基线的性能。
Traditional fine-grained image recognition is required to distinguish different subordinate categories (e.g., birds species) based on the visual cues beneath raw images. Due to both small inter-class variations and large intra-class variations, it is desirable to capture the subtle differences between these sub-categories, which is crucial but challenging for fine-grained recognition. Recently, language modality aggregation has been proved as a successful technique to improve visual recognition in the experience. In this paper, we introduce an end-to-end trainable Progressive Mask Attention (PMA) model for fine-grained recognition by leveraging both visual and language modalities. Our Bi-Modal PMA model can not only stage-by-stage capture the most discriminative part in the visual modality by our mask-based fashion, but also explore the out-of-visual-domain knowledge from the language modality in an interactional alignment paradigm. Specifically, at each stage, a self-attention module is proposed to attend to the key patch from images or text descriptions. Besides, a query-relational module is designed to seize the key words/phrases of texts and further bridge the connection between two modalities. Later, the learned representations of bi-modality from multiple stages are aggregated as the final features for recognition. Our Bi-Modal PMA model only needs raw images and raw text descriptions, without requiring bounding boxes/part annotations in images or key word annotations in texts. By conducting comprehensive experiments on fine-grained benchmark datasets, we demonstrate that the proposed method achieves superior performance over the competing baselines, on either vision and language bi-modality or single visual modality.