Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation

Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation
复制标题

不明视频对象:密集、开放世界分割的基准

DOI:
10.1109/iccv48922.2021.01060
复制
发表时间:
2021
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Du Tran
Du Tran
中科院分区:
--
文献类型:
--
作者:
Weiyao Wang;Matt Feiszli;Heng Wang;Du Tran

文献摘要

被引文献

相似文献

目前最先进的目标检测和分割方法在封闭世界假设下工作得很好。这个封闭世界设置假定在训练和部署期间可以使用对象类别列表。然而,许多现实世界的应用需要检测或分割新对象,即在训练期间从未见过的对象类别。在本文中,我们提出了UVO (Unidentified Video Objects,未识别视频对象),这是一个开放世界视频中与类别无关的对象分割的新基准。除了将焦点转移到开放世界设置之外,UVO显着更大,与DAVIS相比提供大约6倍的视频,与YouTube-VO(I)S相比,每个视频的掩码(实例)注释多7倍。UVO也更具挑战性,因为它包括许多视频拥挤的场景和复杂的背景动作。我们还证明了UVO可以用于其他应用,如对象跟踪和超体素分割。我们相信,UVO是一个多功能的测试平台,为研究人员开发开放世界类别无关的对象分割的新方法,并激发新的研究方向,朝着更全面的视频理解超越分类和检测。
Current state-of-the-art object detection and segmentation methods work well under the closed-world assumption. This closed-world setting assumes that the list of object categories is available during training and deployment. However, many real-world applications require detecting or segmenting novel objects, i.e., object categories never seen during training. In this paper, we present, UVO (Unidentified Video Objects), a new benchmark for openworld class-agnostic object segmentation in videos. Besides shifting the focus to the open-world setup, UVO is significantly larger, providing approximately 6 times more videos compared with DAVIS, and 7 times more mask (instance) annotations per video compared with YouTube-VO(I)S. UVO is also more challenging as it includes many videos with crowded scenes and complex background motions. We also demonstrated that UVO can be used for other applications, such as object tracking and super-voxel segmentation. We believe that UVO is a versatile testbed for researchers to develop novel approaches for open-world class-agnostic object segmentation, and inspires new research directions towards a more comprehensive video understanding beyond classification and detection.