Multiview Cross-supervision for Semantic Segmentation

Multiview Cross-supervision for Semantic Segmentation
复制标题

DOI:
--
复制
发表时间:
2018-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Y. Yao;H. Park
Y. Yao;H. Park
中科院分区:
其他
文献类型:
--
作者:
Y. Yao;H. Park

文献摘要

相似文献

本文提出了一种半监督学习框架,定制的语义分割任务,使用多视图图像流。定制任务的一个关键挑战在于由于禁止手动注释工作的要求,标记数据的可访问性有限。我们假设可以利用通过底层3D几何结构链接的多视图图像流,这可以提供额外的监督信号来训练分割模型。我们制定了一个新的交叉监督方法,使用形状信念转移-在一个图像的分割信念是用来预测,通过对极几何类似于形状从剪影的其他图像。形状置信转移提供了未标记数据的分割的上界和下界,其中随着标记视图的数量增加,其间隙渐近地接近于零。我们整合这一理论,设计一种新的网络,是不可知的摄像机校准,网络模型和语义类别,并绕过次优的3D重建的中间过程。我们通过从现实世界的视觉数据(包括非人类物种和社交视频中感兴趣的主题)中识别每个像素的自定义语义类别来验证该网络,其中获得大规模注释数据是不可行的。
This paper presents a semi-supervised learning framework for a customized semantic segmentation task using multiview image streams. A key challenge of the customized task lies in the limited accessibility of the labeled data due to the requirement of prohibitive manual annotation effort. We hypothesize that it is possible to leverage multiview image streams that are linked through the underlying 3D geometry, which can provide an additional supervisionary signal to train a segmentation model. We formulate a new cross-supervision method using a shape belief transfer---the segmentation belief in one image is used to predict that of the other image through epipolar geometry analogous to shape-from-silhouette. The shape belief transfer provides the upper and lower bounds of the segmentation for the unlabeled data where its gap approaches asymptotically to zero as the number of the labeled views increases. We integrate this theory to design a novel network that is agnostic to camera calibration, network model, and semantic category and bypasses the intermediate process of suboptimal 3D reconstruction. We validate this network by recognizing a customized semantic category per pixel from realworld visual data including non-human species and a subject of interest in social videos where attaining large-scale annotation data is infeasible.