Tracking Through Containers and Occluders in the Wild

Tracking Through Containers and Occluders in the Wild
复制标题

DOI:
10.1109/cvpr52729.2023.01326
复制
发表时间:
2023-05
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Basile Van Hoorick;P. Tokmakov;Simon Stent;Jie Li;Carl Vondrick
Basile Van Hoorick;P. Tokmakov;Simon Stent;Jie Li;Carl Vondrick
中科院分区:
其他
文献类型:
--
作者:
Basile Van Hoorick;P. Tokmakov;Simon Stent;Jie Li;Carl Vondrick

文献摘要

相似文献

在杂乱和动态环境中持续跟踪目标仍然是计算机视觉系统的一个困难的挑战。在本文中,我们介绍了TCOW,一个新的基准和模型的视觉跟踪通过严重的遮挡和遏制。我们建立了一个任务,其目标是,给定一个视频序列,分割目标对象的投影范围,以及周围的容器或遮挡物,只要存在。为了研究这个任务,我们创建了一个混合的合成和注释的真实的数据集,以支持监督学习和结构化评估模型性能的各种形式的任务变化,如移动或嵌套的遏制。我们评估了两个最近的基于transformer的视频模型,发现虽然它们在某些任务变化的设置下能够令人惊讶地跟踪目标,但在我们可以声称跟踪模型已经获得真正的物体持久性概念之前,仍然存在相当大的性能差距。
Tracking objects with persistence in cluttered and dynamic environments remains a difficult challenge for computer vision systems. In this paper, we introduce TCOW, a new benchmark and model for visual tracking through heavy occlusion and containment. We set up a task where the goal is to, given a video sequence, segment both the projected extent of the target object, as well as the surrounding container or occluder whenever one exists. To study this task, we create a mixture of synthetic and annotated real datasets to support both supervised learning and structured evaluation of model performance under various forms of task variation, such as moving or nested containment. We evaluate two recent transformer-based video models and find that while they can be surprisingly capable of tracking targets under certain settings of task variation, there remains a considerable performance gap before we can claim a tracking model to have acquired a true notion of object permanence.