Is It Necessary to Transfer Temporal Knowledge for Domain Adaptive Video Semantic Segmentation?

Is It Necessary to Transfer Temporal Knowledge for Domain Adaptive Video Semantic Segmentation?
复制标题

DOI:
10.1007/978-3-031-19812-0_21
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Xinyi Wu;Zhenyao Wu;J. Wan;Lili Ju;Song Wang
Xinyi Wu;Zhenyao Wu;J. Wan;Lili Ju;Song Wang
中科院分区:
其他
文献类型:
--
作者:
Xinyi Wu;Zhenyao Wu;J. Wan;Lili Ju;Song Wang

文献摘要

相似文献

视频语义分割是计算机视觉中的一项基础和重要任务,通常需要大规模的标记数据来训练深度神经网络模型。为了避免费力的手动标记,最近引入了域自适应视频分割方法,通过将知识从自标记的模拟视频的源域转移到未标记的真实世界视频的目标域。然而,它引出了一个有趣的问题-虽然视频到视频的适应是一个自然的想法,但源数据必须是视频吗?在本文中,我们认为,它是没有必要的时间知识转移,因为在目标域中的视频分割的时间连续性可以估计和执行,而无需参考在源域中的视频。这激发了图像到视频域自适应语义分割(I2 VDA)的新框架,其中源域是一组没有时间信息的图像。在这种情况下,我们通过仅基于空间知识的对抗训练来弥合领域差距,并开发了一种新的时间增强策略,通过该策略,目标领域的时间一致性得到了很好的利用和学习。此外,我们引入了一种新的训练方案,通过利用代理网络实时生成伪标签,这对提高对抗训练的稳定性非常有效。两个合成到真实的场景的实验结果表明,所提出的I2 VDA方法可以实现更好的性能比现有的最先进的视频到视频域的自适应方法的视频语义分割。
Video semantic segmentation is a fundamental and important task in computer vision, and it usually requires large-scale labeled data for training deep neural network models. To avoid laborious manual labeling, domain adaptive video segmentation approaches were recently introduced by transferring the knowledge from the source domain of self-labeled simulated videos to the target domain of unlabeled real-world videos. However, it leads to an interesting question – while video-to-video adaptation is a natural idea,are the source data required to be videos?In this paper, we argue that it is not necessary to transfer temporal knowledge since the temporal continuity of video segmentation in the target domain can be estimated and enforced without reference to videos in the source domain. This motivates a new framework of Image-to-Video Domain Adaptive Semantic Segmentation (I2VDA), where the source domain is a set of images without temporal information. Under this setting, we bridge the domain gap via adversarial training based only on the spatial knowledge, and develop a novel temporal augmentation strategy, through which the temporal consistency in the target domain is well-exploited and learned. In addition, we introduce a new training scheme by leveraging a proxy network to produce pseudo-labels on-the-fly, which is very effective to improve the stability of adversarial training. Experimental results on two synthetic-to-real scenarios show that the proposed I2VDA method can achieve even better performance on video semantic segmentation than existing state-of-the-art video-to-video domain adaption approaches.