Looking for the Devil in the Details: Learning Trilinear Attention Sampling Network for Fine-Grained Image Recognition

Looking for the Devil in the Details: Learning Trilinear Attention Sampling Network for Fine-Grained Image Recognition
复制标题

DOI:
10.1109/cvpr.2019.00515
复制
发表时间:
2019-03
期刊:
2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Heliang Zheng;Jianlong Fu;Zhengjun Zha;Jiebo Luo
Heliang Zheng;Jianlong Fu;Zhengjun Zha;Jiebo Luo
中科院分区:
其他
文献类型:
--
作者:
Heliang Zheng;Jianlong Fu;Zhengjun Zha;Jiebo Luo

文献摘要

被引文献

相似文献

学习微妙但有区别的特征(例如鸟的喙和眼睛)在细粒度图像识别中发挥着重要作用。现有的基于注意力的方法通过定位和放大重要部分来学习细粒度的细节,这通常会受到零件数量有限和计算成本高昂的影响。在本文中,我们建议通过三线性注意力采样网络(TASN)以有效的师生方式从数百个部分提案中学习这种细粒度的特征。具体来说,TASN 由 1)一个三线性注意力模块组成,它通过对通道间关系进行建模来生成注意力图;2)一个基于注意力的采样器,以高分辨率突出显示关注的部分;3)一个特征提取器,它通过权重共享和特征保留策略将部分特征提取为对象级特征。大量实验验证了 TASN 在 iNaturalist-2017、CUB-Bird 和 Stanley-Cars 数据集中使用最具竞争力的方法在相同设置下产生最佳性能。
Learning subtle yet discriminative features (e.g., beak and eyes for a bird) plays a significant role in fine-grained image recognition. Existing attention-based approaches localize and amplify significant parts to learn fine-grained details, which often suffer from a limited number of parts and heavy computational cost. In this paper, we propose to learn such fine-grained features from hundreds of part proposals by Trilinear Attention Sampling Network (TASN) in an efficient teacher-student manner. Specifically, TASN consists of 1) a trilinear attention module, which generates attention maps by modeling the inter-channel relationships, 2) an attention-based sampler which highlights attended parts with high resolution, and 3) a feature distiller, which distills part features into an object-level feature by weight sharing and feature preserving strategies. Extensive experiments verify that TASN yields the best performance under the same settings with the most competitive approaches, in iNaturalist-2017, CUB-Bird, and Stanford-Cars datasets.