Weakly-supervised Object Representation Learning for Few-shot Semantic Segmentation

Weakly-supervised Object Representation Learning for Few-shot Semantic Segmentation
复制标题

DOI:
10.1109/wacv48630.2021.00154
复制
发表时间:
2021-01
期刊:
2021 IEEE Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Xiaowen Ying;Xin Li;M. Chuah
Xiaowen Ying;Xin Li;M. Chuah
中科院分区:
其他
文献类型:
--
作者:
Xiaowen Ying;Xin Li;M. Chuah

文献摘要

相似文献

训练语义分割模型需要大量密集注释的图像数据集,这些数据集的获取成本很高。一旦训练完成,也很难向这样的分割模型添加新的对象类别。在本文中,我们解决了少镜头语义分割问题,其目的是执行图像分割任务上看不见的对象类别仅仅基于一个或几个支持的例子(S)。解决这一少镜头分割问题的关键在于有效地利用支持样本中的对象信息将查询图像中的目标对象从背景中分离出来。虽然现有的方法通常通过平均支持图像中的局部特征来生成对象级表示,但我们证明了这种对象表示通常是嘈杂的,并且不那么区分。为了解决这个问题,我们设计了一个对象表示生成器(ORG)模块,它可以有效地从支持图像中聚合局部对象特征,生成更好的对象级表示。ORG模块可以嵌入到网络中,并以弱监督的方式进行端到端的训练,而无需额外的人工注释。我们将此设计到一个修改后的编码器-解码器网络,提出了一个强大而有效的框架,少数镜头语义分割。在Pascal-VOC和MS-COCO数据集上的实验结果表明,与现有方法相比,该方法在单次和五次设置下都具有更好的性能。
Training a semantic segmentation model requires large densely-annotated image datasets that are costly to obtain. Once the training is done, it is also difficult to add new object categories to such segmentation models. In this paper, we tackle the few-shot semantic segmentation problem, which aims to perform image segmentation task on unseen object categories merely based on one or a few support example(s). The key to solving this few-shot segmentation problem lies in effectively utilizing object information from support examples to separate target objects from the background in a query image. While existing methods typically generate object-level representations by averaging local features in support images, we demonstrate that such object representations are typically noisy and less distinguishing. To solve this problem, we design an object representation generator (ORG) module which can effectively aggregate local object features from support im- age(s) and produce better object-level representation. The ORG module can be embedded into the network and trained end-to-end in a weakly-supervised fashion without extra human annotation. We incorporate this design into a modified encoder-decoder network to present a powerful and efficient framework for few-shot semantic segmentation. Experimental results on the Pascal-VOC and MS-COCO datasets show that our approach achieves better performance compared to existing methods under both one-shot and five-shot settings.