Recurrently exploring class-wise attention in a hybrid convolutional and bidirectional LSTM network for multi-label aerial image classification

Recurrently exploring class-wise attention in a hybrid convolutional and bidirectional LSTM network for multi-label aerial image classification
复制标题

DOI:
10.1016/j.isprsjprs.2019.01.015
复制
发表时间:
2019-03-01
影响因子:
12.7
通讯作者:
Zhu, Xiao Xiang
Zhu, Xiao Xiang
中科院分区:
工程技术1区
文献类型:
--
作者:
Hua, Yuansheng;Mou, Lichao;Zhu, Xiao Xiang

文献摘要

被引文献

相似文献

航空图像分类在遥感界具有重要意义,过去几年已经开展了许多研究。在这些研究中,大多数都专注于将图像分类为一个语义标签,而在现实世界中,航拍图像通常与多个标签相关联,例如在我们的例子中,多个对象级标签。此外,给定高分辨率航空图像中当前物体的全面图片可以提供对研究区域更深入的了解。由于这些原因,航空图像多标签分类已引起越来越多的关注。然而,社区中现有方法的一个共同限制是,各种类的共现关系,即所谓的类依赖性,没有得到充分探索,并导致不周全的决策。在本文中,我们针对此任务提出了一种新颖的端到端网络,即基于类别注意的卷积和双向 LSTM 网络(CA-Cony-BiLSTM)。该网络由三个不可或缺的组件组成:(1)特征提取模块,(2)类别注意力学习层,(3)基于 LSTM 的双向子网络。特别是,特征提取模块旨在提取细粒度的语义特征图,而类别注意力学习层旨在捕获有区别的类别特定特征。作为最重要的部分,基于 LSTM 的双向子网络对两个方向的底层类依赖关系进行建模,并生成结构化的多个对象标签。 UCM多标签数据集和DFC15多标签数据集上的实验结果定量和定性地验证了我们模型的有效性。
Aerial image classification is of great significance in the remote sensing community, and many researches have been conducted over the past few years. Among these studies, most of them focus on categorizing an image into one semantic label, while in the real world, an aerial image is often associated with multiple labels, e.g., multiple object-level labels in our case. Besides, a comprehensive picture of present objects in a given high-resolution aerial image can provide a more in-depth understanding of the studied region. For these reasons, aerial image multi-label classification has been attracting increasing attention. However, one common limitation shared by existing methods in the community is that the co-occurrence relationship of various classes, so-called class dependency, is underexplored and leads to an inconsiderate decision. In this paper, we propose a novel end-to end network, namely class-wise attention-based convolutional and bidirectional LSTM network (CA-Cony-BiLSTM), for this task. The proposed network consists of three indispensable components: (1) a feature extraction module, (2) a class attention learning layer, and (3) a bidirectional LSTM-based sub-network. Particularly, the feature extraction module is designed for extracting fine-grained semantic feature maps, while the class attention learning layer aims at capturing discriminative class-specific features. As the most important part, the bidirectional LSTM-based sub-network models the underlying class dependency in both directions and produce structured multiple object labels. Experimental results on UCM multi-label dataset and DFC15 multi label dataset validate the effectiveness of our model quantitatively and qualitatively.