Self-Supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

Self-Supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation
复制标题

DOI:
10.1109/cvpr42600.2020.01229
复制
发表时间:
2020-04
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yude Wang;Jie Zhang;Meina Kan;S. Shan;Xilin Chen
Yude Wang;Jie Zhang;Meina Kan;S. Shan;Xilin Chen
中科院分区:
其他
文献类型:
--
作者:
Yude Wang;Jie Zhang;Meina Kan;S. Shan;Xilin Chen

文献摘要

被引文献

相似文献

图像级弱监督语义分割是一个具有挑战性的问题,近年来得到了深入的研究。大多数高级解决方案利用类激活图(CAM)。然而,由于全监督和弱监督之间存在差距,CAM很难起到对象掩模的作用。在本文中,我们提出了一个自我监督的同变注意机制(SEAM),发现额外的监督和缩小差距。我们的方法是基于这样的观察,即等方差是全监督语义分割中的一个隐式约束,其像素级标签在数据增强过程中与输入图像进行相同的空间变换。然而,这种约束在通过图像级监督训练的CAM上丢失。因此,我们提出了一致性正则化预测CAM从各种变换的图像,为网络学习提供自我监督。此外,我们提出了一个像素相关性模块(PCM),它利用上下文外观信息和细化预测当前像素的相似邻居,导致进一步提高CAM一致性。在PASCAL VOC 2012数据集上进行的大量实验表明,我们的方法在使用相同监督级别的情况下优于最先进的方法。代码已在线发布。
Image-level weakly supervised semantic segmentation is a challenging problem that has been deeply studied in recent years. Most of advanced solutions exploit class activation map (CAM). However, CAMs can hardly serve as the object mask due to the gap between full and weak supervisions. In this paper, we propose a self-supervised equivariant attention mechanism (SEAM) to discover additional supervision and narrow the gap. Our method is based on the observation that equivariance is an implicit constraint in fully supervised semantic segmentation, whose pixel-level labels take the same spatial transformation as the input images during data augmentation. However, this constraint is lost on the CAMs trained by image-level supervision. Therefore, we propose consistency regularization on predicted CAMs from various transformed images to provide self-supervision for network learning. Moreover, we propose a pixel correlation module (PCM), which exploits context appearance information and refines the prediction of current pixel by its similar neighbors, leading to further improvement on CAMs consistency. Extensive experiments on PASCAL VOC 2012 dataset demonstrate our method outperforms state-of-the-art methods using the same level of supervision. The code is released online.