Context-Aware Relative Object Queries to Unify Video Instance and Panoptic Segmentation

Context-Aware Relative Object Queries to Unify Video Instance and Panoptic Segmentation
复制标题

DOI:
10.1109/cvpr52729.2023.00617
复制
发表时间:
2023-06
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Anwesa Choudhuri;Girish V. Chowdhary;A. Schwing
Anwesa Choudhuri;Girish V. Chowdhary;A. Schwing
中科院分区:
其他
文献类型:
--
作者:
Anwesa Choudhuri;Girish V. Chowdhary;A. Schwing

文献摘要

相似文献

对象查询已经成为一个强大的抽象,一般表示对象的建议。然而,它们用于诸如视频分割之类的时间任务提出了两个问题:1)如何顺序地处理帧并跨帧无缝地传播对象查询。每帧使用独立的对象查询不允许跟踪,并且需要后处理。2)如何产生时间上一致的,但表达对象的查询模型的外观和位置的变化。一次使用整个视频不会捕捉位置变化,也不会扩展到长视频。作为这两个问题的答案之一,我们提出了“上下文感知的相对对象查询”,这是不断传播的逐帧。它们无缝地跟踪对象并处理对象的遮挡和再现,而无需后处理。此外,我们发现上下文感知的相对对象查询更好地捕捉运动中的对象的位置变化。我们评估所提出的方法在三个具有挑战性的任务:视频实例分割,多对象跟踪和分割,视频全景分割。使用相同的方法和架构,我们在具有挑战性的OVIS、Youtube-VIS、Cityscapes-VPS、MOTS 2020和KITTI-MOTS数据上匹配或超越了最先进的结果。
Object queries have emerged as a powerful abstraction to generically represent object proposals. However, their use for temporal tasks like video segmentation poses two questions: 1) How to process frames sequentially and propagate object queries seamlessly across frames. Using independent object queries per frame doesn't permit tracking, and requires post-processing. 2) How to produce temporally consistent, yet expressive object queries that model both appearance and position changes. Using the entire video at once doesn't capture position changes and doesn't scale to long videos. As one answer to both questions we propose 'context-aware relative object queries', which are continuously propagated frame-by-frame. They seamlessly track objects and deal with occlusion and re-appearance of objects, without post-processing. Further, we find context-aware relative object queries better capture position changes of objects in motion. We evaluate the proposed approach across three challenging tasks: video instance segmentation, multi-object tracking and segmentation, and video panoptic segmentation. Using the same approach and architecture, we match or surpass state-of-the art results on the diverse and challenging OVIS, Youtube-VIS, Cityscapes-VPS, MOTS 2020 and KITTI-MOTS data.