Video K-Net: A Simple, Strong, and Unified Baseline for Video Segmentation

Video K-Net: A Simple, Strong, and Unified Baseline for Video Segmentation
复制标题

DOI:
10.1109/cvpr52688.2022.01828
复制
发表时间:
2022-04
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Xiangtai Li;Wenwei Zhang;Jiangmiao Pang;Kai Chen;Guangliang Cheng;Yunhai Tong;Chen Change Loy
Xiangtai Li;Wenwei Zhang;Jiangmiao Pang;Kai Chen;Guangliang Cheng;Yunhai Tong;Chen Change Loy
中科院分区:
其他
文献类型:
--
作者:
Xiangtai Li;Wenwei Zhang;Jiangmiao Pang;Kai Chen;Guangliang Cheng;Yunhai Tong;Chen Change Loy

文献摘要

被引文献

相似文献

本文介绍了视频K-Net,一个简单,强大,统一的框架,完全端到端的视频全景分割。该方法是建立在K-Net,通过一组可学习的内核统一图像分割的方法。我们观察到,这些来自K-Net的可学习内核(对对象外观和上下文进行编码)可以自然地将视频帧中的相同实例关联起来。受到这一观察的启发,视频K-Net学会了通过简单的基于内核的外观建模和跨时间内核交互来同时分割和跟踪视频中的“事物”和“东西”。尽管简单,它实现了最先进的视频全景分割结果的城市景观VPS和KITTI-STEP没有花里胡哨。特别是在KITTI-STEP上,简单的方法可以比以前的方法提高近12%的相对改进。我们还验证了它在视频语义分割上的泛化,在VSPW数据集上,我们将各种基线提高了2%。此外,我们将K-Net扩展到剪辑级视频框架中,用于视频实例分割,在YouTube-2019验证集上,ResNet 50骨干获得40.5%,Swin-base获得51.5%的mAP。我们希望这种简单而有效的方法可以作为视频分割的新的灵活基线。11代码和模型都在这里发布。
This paper presents Video K-Net, a simple, strong, and unified framework for fully end-to-end video panoptic seg-mentation. The method is built upon K-Net, a method that unifies image segmentation via a group of learnable ker-nels. We observe that these learnable kernels from K-Net, which encode object appearances and contexts, can naturally associate identical instances across video frames. Motivated by this observation, Video K-Net learns to simultaneously segment and track “things” and “stuff” in a video with simple kernel-based appearance modeling and cross-temporal kernel interaction. Despite the simplicity, it achieves state-of-the-art video panoptic segmentation results on Citscapes-VPS and KITTI-STEP without bells and whistles. In particular on KITTI-STEP, the simple method can boost almost 12% relative improvements over previous methods. We also validate its generalization on video semantic segmentation, where we boost various baselines by 2% on the VSPW dataset. Moreover, we extend K-Net into clip-level video framework for video instance segmentation where we obtain 40.5% for ResNet50 backbone and 51.5% mAP for Swin-base on YouTube-2019 validation set. We hope this simple yet effective method can serve as a new flexible baseline in video segmentation.11Both code and models are released at here.