Contrastive Learning for Label Efficient Semantic Segmentation

Contrastive Learning for Label Efficient Semantic Segmentation
复制标题

DOI:
10.1109/iccv48922.2021.01045
复制
发表时间:
2020-12
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Xiangyu Zhao;Raviteja Vemulapalli;P. A. Mansfield;Boqing Gong;Bradley Green;Lior Shapira;Ying Wu
Xiangyu Zhao;Raviteja Vemulapalli;P. A. Mansfield;Boqing Gong;Bradley Green;Lior Shapira;Ying Wu
中科院分区:
其他
文献类型:
--
作者:
Xiangyu Zhao;Raviteja Vemulapalli;P. A. Mansfield;Boqing Gong;Bradley Green;Lior Shapira;Ying Wu

文献摘要

被引文献

相似文献

收集用于语义分割任务的标记数据既昂贵又耗时,因为它需要密集的像素级注释。虽然最近基于卷积神经网络(CNN)的语义分割方法通过使用大量的标记训练数据取得了令人印象深刻的结果,但它们的性能随着标记数据量的减少而显著下降。这是因为用事实上的交叉熵损失训练的深层CNN很容易过度适应少量的标记数据。为了解决这个问题,我们提出了一种简单而有效的基于对比学习的训练策略,在该策略中,我们首先使用基于像素的、基于标签的对比损失来预训练网络,然后使用交叉熵损失来微调网络。这种方法增加了类内紧凑性和类间可分离性,从而得到了更好的像素分类器。我们使用CitySees和Pascal VOC 2012分割数据集验证了所提出的训练策略的有效性。我们的结果表明,在标签数据量有限的情况下,具有所提出的对比损失的预训练会导致很大的性能提升(在某些设置中绝对改善超过20%)。在许多情况下,提出的不使用任何额外数据的对比性预训练策略能够与广泛使用的使用100多万个额外标记图像的ImageNet预训练策略相匹配或优于后者。
Collecting labeled data for the task of semantic segmentation is expensive and time-consuming, as it requires dense pixel-level annotations. While recent Convolutional Neural Network (CNN) based semantic segmentation approaches have achieved impressive results by using large amounts of labeled training data, their performance drops significantly as the amount of labeled data decreases. This happens because deep CNNs trained with the de facto cross-entropy loss can easily overfit to small amounts of labeled data. To address this issue, we propose a simple and effective contrastive learning-based training strategy in which we first pretrain the network using a pixel-wise, label-based contrastive loss, and then fine-tune it using the cross-entropy loss. This approach increases intra-class compactness and inter-class separability, thereby resulting in a better pixel classifier. We demonstrate the effectiveness of the proposed training strategy using the Cityscapes and PASCAL VOC 2012 segmentation datasets. Our results show that pretraining with the proposed contrastive loss results in large performance gains (more than 20% absolute improvement in some settings) when the amount of labeled data is limited. In many settings, the proposed contrastive pretraining strategy, which does not use any additional data, is able to match or outperform the widely-used ImageNet pretraining strategy that uses more than a million additional labeled images.