A Multi-Attention UNet for Semantic Segmentation in Remote Sensing Images

A Multi-Attention UNet for Semantic Segmentation in Remote Sensing Images
复制标题

DOI:
10.3390/sym14050906
复制
发表时间:
2022-05-01
期刊:
影响因子:
2.7
通讯作者:
Feng, Suting
Feng, Suting
中科院分区:
综合性期刊4区
文献类型:
--
作者:
Sun, Yu;Bi, Fukun;Feng, Suting

文献摘要

被引文献

相似文献

近年来,随着深度学习的发展,遥感图像语义分割逐渐成为计算机视觉领域的热点问题。然而,多类别目标的分割仍然是一个难题。为了解决不同类别中精度差和多尺度的问题,我们提出了一种基于多重注意力的 UNet(MA-UNet)。具体来说,我们提出了一种基于简单注意模块的残差编码器,以提高主干网对细粒度特征的提取能力。通过对最低层特征使用多头自注意力,重建给定特征图的语义表示,进一步实现对不同类别像素的细粒度分割。然后,针对不同类别多尺度的问题,我们增加下采样的数量来细分不同尺度下目标的特征尺寸,并在不同特征融合阶段使用通道注意力和空间注意力,更好地融合不同尺度下目标的特征信息。我们在 WHDLD 数据集和 DLRSD 数据集上进行了实验。结果表明,通过多个视觉注意特征增强,我们的方法在 WHDLD 数据集上实现了 63.94% 的平均交并集 (IOU);这个结果比 UNet 高 4.27%,并且在 DLRSD 数据集上,我们的方法的平均 IOU 将 UNet 的 56.17% 提高到 61.90%,同时超过了其他先进方法。
In recent years, with the development of deep learning, semantic segmentation for remote sensing images has gradually become a hot issue in computer vision. However, segmentation for multicategory targets is still a difficult problem. To address the issues regarding poor precision and multiple scales in different categories, we propose a UNet, based on multi-attention (MA-UNet). Specifically, we propose a residual encoder, based on a simple attention module, to improve the extraction capability of the backbone for fine-grained features. By using multi-head self-attention for the lowest level feature, the semantic representation of the given feature map is reconstructed, further implementing fine-grained segmentation for different categories of pixels. Then, to address the problem of multiple scales in different categories, we increase the number of down-sampling to subdivide the feature sizes of the target at different scales, and use channel attention and spatial attention in different feature fusion stages, to better fuse the feature information of the target at different scales. We conducted experiments on the WHDLD datasets and DLRSD datasets. The results show that, with multiple visual attention feature enhancements, our method achieves 63.94% mean intersection over union (IOU) on the WHDLD datasets; this result is 4.27% higher than that of UNet, and on the DLRSD datasets, the mean IOU of our methods improves UNet's 56.17% to 61.90%, while exceeding those of other advanced methods.