Scale-aware attention network for weakly supervised semantic segmentation

Scale-aware attention network for weakly supervised semantic segmentation
复制标题

DOI:
10.1016/j.neucom.2022.04.006
复制
发表时间:
2022-04
期刊:
影响因子:
6
通讯作者:
Zhiyuan Cao;Yufei Gao;Jia-cai Zhang
Zhiyuan Cao;Yufei Gao;Jia-cai Zhang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhiyuan Cao;Yufei Gao;Jia-cai Zhang

文献摘要

相似文献

使用图像级标签的弱监督语义分割(WSSS)大大减轻了获得大量像素级注释的负担。为了创建伪分割标签,大多数WSSS算法依赖于类激活图(CAM)中的局部响应区域。然而,这样的激活图只关注对象的局部判别部分,因为分类网络不需要整个对象来优化目标函数。为了提高网络的能力,更多地集中在非歧视性的部分的对象,并产生高质量的伪面具,尺度感知的注意力网络(SAN)的建议。具体地说,金字塔的注意力模块被引入到传播的歧视性信息相邻的对象区域,通过自适应地选择上下文特征的卷积金字塔与不同的过滤器尺度。为了更好地利用定位图在不同尺度下的互补信息,提出了一种带联合损失的多尺度预测融合结构。在推理阶段,通过对多尺度预测进行加权融合,获得密集且完整的定位图,然后用于训练分割模型。这种新型SAN在一系列实验中证明了其有效性。在PASCAL VOC 2012分割测试集上,与相同监督水平下的其他方法相比,它实现了71.9% mIoU的最新结果。
Weakly supervised semantic segmentation (WSSS) using image-level labels greatly alleviates the burden of obtaining large amounts of pixel-wise annotations. To create pseudo segmentation labels, most WSSS algorithms rely on the local response regions in the class activation maps (CAMs). However, such activation maps only focus on the local discriminative parts of the object, because the classification network does not require the entire object to optimize the objective function. To enhance the network’s ability to focus more on non-discriminative parts of the object and generate high-quality pseudo-masks, the Scale-aware Attention Network (SAN) is proposed. Specifically, a pyramidal attention module is introduced to propagate discriminative information to adjacent object regions by adaptively selecting contextual features from the convolutional pyramid with varied filter scales. A multi-scale prediction fusion structure with a joint loss is proposed to make better use of the complementary information of localization maps in different scales. The dense and integral localization maps are obtained in the inference stage by weighted fusion of the multi-scale predictions, which are then used to train segmentation models. This novel SAN has demonstrated its effectiveness in a series of experiments. It has achieved a state-of-the-art result of 71.9% mIoU on the PASCAL VOC 2012 segmentation test set compared with other approaches under the same level of supervision.