End-to-end weakly supervised semantic segmentation with reliable region mining

End-to-end weakly supervised semantic segmentation with reliable region mining
复制标题

具有可靠区域挖掘的端到端弱监督语义分割

DOI:
10.1016/j.patcog.2022.108663
复制
发表时间:
2022
影响因子:
8
通讯作者:
Yao Zhao
Yao Zhao
中科院分区:
计算机科学1区
文献类型:
--
作者:
Bingfeng Zhang;Jimin Xiao;Yunchao Wei;Kaizhu Huang;Shan Luo;Yao Zhao

文献摘要

相似文献

弱监督语义分割是一个具有挑战性的任务,它只需要图像级的标签作为监督,但产生用于测试的像素级预测。为了解决这样一个具有挑战性的任务,大多数当前的方法首先生成伪像素掩码,然后将其馈送到单独的语义分割网络中。然而,这些两步方法具有高度复杂性,并且难以整体训练。在这项工作中,我们利用图像级标签来生成可靠的像素级注释,并设计一个完全端到端的网络来学习预测分割图。具体地说,我们首先利用图像分类分支来生成注释类别的类别激活图,这些类别激活图被进一步修剪成微小的可靠对象/背景区域。这样的可靠区域然后被直接用作分割分支的地面实况标签,其中全局信息和局部信息子分支两者被用于生成准确的像素级预测。此外,提出了一种新的联合损失,同时考虑浅层和高级功能。尽管其表面上的简单性,我们的端到端解决方案实现了具有竞争力的mIoU分数(瓦尔:65.4%,测试:65.3%)的Pascal VOC相比,两步同行。通过将我们的一步方法扩展为两步方法,我们在Pascal VOC 2012数据集上获得了新的最先进的性能(瓦尔:69.3%,test:69.2%)。代码可从以下网址获得:https://github.com/zbf1991/RRM。
Weakly supervised semantic segmentation is a challenging task that only takes image-level labels as supervision but produces pixel-level predictions for testing. To address such a challenging task, most current approaches generate pseudo pixel masks first that are then fed into a separate semantic segmentation network. However, these two-step approaches suffer from high complexity and being hard to train as a whole. In this work, we harness the image-level labels to produce reliable pixel-level annotations and design a fully end-to-end network to learn to predict segmentation maps. Concretely, we firstly leverage an image classification branch to generate class activation maps for the annotated categories, which are further pruned into tiny reliable object/background regions. Such reliable regions are then directly served as ground-truth labels for the segmentation branch, where both global information and local information sub-branches are used to generate accurate pixel-level predictions. Furthermore, a new joint loss is proposed that considers both shallow and high-level features. Despite its apparent simplicity, our end-to-end solution achieves competitive mIoU scores (val: 65.4%,test: 65.3%) on Pascal VOC compared with the two-step counterparts. By extending our one-step method to two-step, we get a new state-of-the-art performance on the Pascal VOC 2012 dataset(val: 69.3%,test: 69.2%). Code is available at: https://github.com/zbf1991/RRM.