Multi-Label Visual Feature Learning with Attentional Aggregation

Multi-Label Visual Feature Learning with Attentional Aggregation
复制标题

DOI:
10.1109/wacv45572.2020.9093311
复制
发表时间:
2020-03
期刊:
2020 IEEE Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Ziqiao Guan;K. Yager;Dantong Yu;Hong Qin
Ziqiao Guan;K. Yager;Dantong Yu;Hong Qin
中科院分区:
其他
文献类型:
--
作者:
Ziqiao Guan;K. Yager;Dantong Yu;Hong Qin

文献摘要

相似文献

如今,卷积神经网络 (CNN) 已经扩展到科学界的专门应用,否则这些应用将无法得到充分解决。在本文中,我们系统地研究了材料科学中X射线散射图像的多标签标注问题。对于此应用,我们通过训练 CNN 解决了​​一个公开挑战——识别具有扩散背景干扰的弱散射图案,这在科学成像中很常见。我们阐明了注意力聚合模块(AAM)来增强特征表示。首先,我们使用数据驱动的注意力图重新加权并突出显示图像中的重要特征。我们将注意力图分解为通道和空间注意力组件。在空间注意力组件中,我们设计了一种机制来生成适合多样化多标签学习的多个空间注意力图。然后,我们通过执行特征聚合将增强的局部特征压缩为非局部表示。注意力和聚合都被设计为具有可学习参数的网络层,以便 CNN 训练保持流畅的端到端,并且我们将其在网络中应用几次,以便特征增强是多尺度的。我们对 CNN 训练和测试以及迁移学习进行了广泛的实验,实证研究证实我们的方法增强了科学成像视觉特征的判别力。
Today convolutional neural networks (CNNs) have reached out to specialized applications in science communities that otherwise would not be adequately tackled. In this paper, we systematically study a multi-label annotation problem of x-ray scattering images in material science. For this application, we tackle an open challenge with training CNNs — identifying weak scattered patterns with diffuse background interference, which is common in scientific imaging. We articulate an Attentional Aggregation Module (AAM) to enhance feature representations. First, we reweight and highlight important features in the images using data-driven attention maps. We decompose the attention maps into channel and spatial attention components. In the spatial attention component, we design a mechanism to generate multiple spatial attention maps tailored for diversified multi-label learning. Then, we condense the enhanced local features into non-local representations by performing feature aggregation. Both attention and aggregation are designed as network layers with learnable parameters so that CNN training remains fluidly end-to-end, and we apply it in-network a few times so that the feature enhancement is multi-scale. We conduct extensive experiments on CNN training and testing, as well as transfer learning, and empirical studies confirm that our method enhances the discriminative power of visual features of scientific imaging.