Deformable attention-oriented feature pyramid network for semantic segmentation

Deformable attention-oriented feature pyramid network for semantic segmentation
复制标题

DOI:
10.1016/j.knosys.2022.109623
复制
发表时间:
2022-08
期刊:
Knowl. Based Syst.
影响因子:
--
通讯作者:
Lei Lu;Yun Xiao;Xiaojun Chang;Xuanhong Wang;Pengzhen Ren;Zhe Ren
Lei Lu;Yun Xiao;Xiaojun Chang;Xuanhong Wang;Pengzhen Ren;Zhe Ren
中科院分区:
其他
文献类型:
--
作者:
Lei Lu;Yun Xiao;Xiaojun Chang;Xuanhong Wang;Pengzhen Ren;Zhe Ren

文献摘要

相似文献

在计算机视觉领域,金字塔特征的使用可以显著提高网络性能。然而,语义信息的不对齐和小尺度特征的尺度限制导致特征贡献的不平衡,严重限制了特征金字塔网络的性能。为了解决特征贡献不平衡导致的模型效率下降问题,本文提出了一种可变形的面向注意力的特征金字塔网络(DAFPN)。与以往的模型只关注特征之间的语义信息不同,DAFPN使用可变形的注意机制对多个特征之间的关系进行建模,然后在金字塔特征融合过程中对它们进行合并。在DAFPN的基础上,我们进一步提出了一种基于全变换的语义分割头,实现了高性能和良好的可扩展性。在多个主干网上的比较表明,我们提出的模型优于基线模型。在相同的条件下,我们的方法可以将mIoU提高1 ~ 4%,高于基线语义分割模型。
In the field of computer vision, the use of pyramid features can significantly improve network performance. However, the misalignment of semantic information and the scale limitation of small-scale features lead to an imbalance of feature contributions, which severely limits the performance of the feature pyramid network. In order to solve the problem of model efficiency decline caused by feature contribution imbalance, in this paper, we propose a deformable attention-oriented feature pyramid network (DAFPN). Unlike previous models, which focus solely on the semantic information between features, DAFPN uses the deformable attention mechanism to model the relationship between multiple features and then merges them in the pyramid feature fusion process. Based on DAFPN, we further propose a fully transformer-based semantic segmentation head, which achieves high performance and good scalability. Comparisons on multiple backbones reveal that our proposed model outperforms the baseline model. Under the same conditions, our method can improve the mIoU by 1∼ 4%, which is higher than the baseline semantic segmentation model.