Contextual Guided Segmentation Framework for Semi-supervised Video Instance Segmentation

Contextual Guided Segmentation Framework for Semi-supervised Video Instance Segmentation
复制标题

DOI:
10.1007/s00138-022-01278-x
复制
发表时间:
2021-06
影响因子:
3.3
通讯作者:
Trung-Nghia Le;Tam V. Nguyen;M. Tran
Trung-Nghia Le;Tam V. Nguyen;M. Tran
中科院分区:
计算机科学4区
文献类型:
--
作者:
Trung-Nghia Le;Tam V. Nguyen;M. Tran

文献摘要

相似文献

在本文中,我们提出了上下文引导分割(CGS)框架的视频实例分割三遍。在第一遍中,即,预览分割,我们提出实例重新识别流程来估计每个实例的主要属性(即,人类/非人类、刚性/可变形、已知/未知类别)。在第二遍中,即,上下文分割,我们介绍多个上下文分割方案。对于人类的例子,我们开发了一个帧中的图像引导分割沿着对象流,以纠正和完善跨帧的结果。对于非人类实例,如果实例在外观上有很大的变化并且属于已知类别(可以从初始掩码中推断),则采用实例分割。如果非人类实例几乎是刚性的,则我们在来自视频序列的第一帧的合成图像上训练FCN。在最后一遍中,即,引导分割,我们开发了一种新的细粒度的非矩形感兴趣区域(ROI)的分割方法。自然形状的ROI通过应用来自当前帧的相邻帧的引导注意来生成,以减少不同重叠实例的分割中的模糊性。前向掩码传播之后是后向掩码传播,以进一步恢复由于重新出现的实例、快速运动、遮挡或严重变形而丢失的实例片段。最后,每个帧中的实例根据它们的深度值、人类和非人类对象交互以及稀有实例优先级进行合并。在DAVIS测试挑战数据集上进行的实验证明了我们提出的框架的有效性。在2017-2019年DAVIS挑战赛中,我们在全球得分、区域相似性和轮廓准确性方面分别以75.4%、72.4%和78.4%的成绩获得了第三名。
In this paper, we propose contextual guided segmentation (CGS) framework for video instance segmentation in three passes. In the first pass,i.e.,preview segmentation, we propose Instance Re-Identification Flow to estimate main properties of each instance (i.e., human/non-human, rigid/deformable, known/unknown category) by propagating its preview mask to other frames. In the second pass,i.e.,contextual segmentation, we introduce multiple contextual segmentation schemes. For human instance, we develop skeleton-guided segmentation in a frame along with object flow to correct and refine the result across frames. For non-human instance, if the instance has a wide variation in appearance and belongs to known categories (which can be inferred from the initial mask), we adopt instance segmentation. If the non-human instance is nearly rigid, we train FCNs on synthesized images from the first frame of a video sequence. In the final pass,i.e.,guided segmentation, we develop a novel fined-grained segmentation method on non-rectangular regions of interest (ROIs). The natural-shaped ROI is generated by applying guided attention from the neighbor frames of the current one to reduce the ambiguity in the segmentation of different overlapping instances. Forward mask propagation is followed by backward mask propagation to further restore missing instance fragments due to re-appeared instances, fast motion, occlusion, or heavy deformation. Finally, instances in each frame are merged based on their depth values, together with human and non-human object interaction and rare instance priority. Experiments conducted on the DAVIS Test-Challenge dataset demonstrate the effectiveness of our proposed framework. We achieved the 3rd consistently in the DAVIS Challenges 2017–2019 with 75.4%, 72.4%, and 78.4% in terms of global score, region similarity, and contour accuracy, respectively.