MAFF-HRNet: Multi-Attention Feature Fusion HRNet for Building Segmentation in Remote Sensing Images

MAFF-HRNet: Multi-Attention Feature Fusion HRNet for Building Segmentation in Remote Sensing Images
复制标题

DOI:
10.3390/rs15051382
复制
发表时间:
2023-02
期刊:
Remote. Sens.
影响因子:
--
通讯作者:
Zhihao Che;Li Shen;Lianzhi Huo;Changmiao Hu;Yanping Wang;Yao Lu;Fukun Bi
Zhihao Che;Li Shen;Lianzhi Huo;Changmiao Hu;Yanping Wang;Yao Lu;Fukun Bi
中科院分区:
其他
文献类型:
--
作者:
Zhihao Che;Li Shen;Lianzhi Huo;Changmiao Hu;Yanping Wang;Yao Lu;Fukun Bi

文献摘要

相似文献

建成区和建筑物是遥感研究的两个主要目标;因此,建成区和建筑物的自动提取受到了广泛的关注。由于边界模糊、对象遮挡和类内不一致,该任务通常很困难。在本文中,我们提出了多注意力特征融合HRNet,MAFF-HRNet,它可以保留更详细的特征以实现准确的语义分割。金字塔特征注意力(PFA)层次结构的设计增强了模型的多级语义表示。此外,我们开发了混合卷积注意(MCA)块,它增加了感受野的捕获范围并克服了类内不一致的问题。为了减轻遮挡造成的干扰,还提出了多尺度注意特征聚合(MAFA)块来增强最终预测图的恢复。我们的方法在 WHU(武汉大学)建筑数据集和马萨诸塞州建筑数据集上进行了系统测试。与其他先进的语义分割模型相比,我们的模型取得了最好的 IoU 结果,分别为 91.69% 和 68.32%。为了进一步评估所提出模型的应用意义,我们将基于World-Cover数据集训练的预训练模型迁移到高分16 m数据集进行测试。定量和定性实验表明,我们的模型可以从遥感图像中准确地分割建筑物和建成区。
Built-up areas and buildings are two main targets in remote sensing research; consequently, automatic extraction of built-up areas and buildings has attracted extensive attention. This task is usually difficult because of boundary blur, object occlusion, and intra-class inconsistency. In this paper, we propose the multi-attention feature fusion HRNet, MAFF-HRNet, which can retain more detailed features to achieve accurate semantic segmentation. The design of a pyramidal feature attention (PFA) hierarchy enhances the multilevel semantic representation of the model. In addition, we develop a mixed convolutional attention (MCA) block, which increases the capture range of receptive fields and overcomes the problem of intra-class inconsistency. To alleviate interference due to occlusion, a multiscale attention feature aggregation (MAFA) block is also proposed to enhance the restoration of the final prediction map. Our approach was systematically tested on the WHU (Wuhan University) Building Dataset and the Massachusetts Buildings Dataset. Compared with other advanced semantic segmentation models, our model achieved the best IoU results of 91.69% and 68.32%, respectively. To further evaluate the application significance of the proposed model, we migrated a pretrained model based on the World-Cover Dataset training to the Gaofen 16 m dataset for testing. Quantitative and qualitative experiments show that our model can accurately segment buildings and built-up areas from remote sensing images.