Learning to Track Instances without Video Annotations

Learning to Track Instances without Video Annotations
复制标题

DOI:
10.1109/cvpr46437.2021.00857
复制
发表时间:
2021-04
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yang Fu;Sifei Liu;Umar Iqbal-;Shalini De Mello;Humphrey Shi;J. Kautz
Yang Fu;Sifei Liu;Umar Iqbal-;Shalini De Mello;Humphrey Shi;J. Kautz
中科院分区:
其他
文献类型:
--
作者:
Yang Fu;Sifei Liu;Umar Iqbal-;Shalini De Mello;Humphrey Shi;J. Kautz

文献摘要

被引文献

相似文献

多个实例的跟踪分割掩码已经得到了深入的研究,但仍然面临着两个基本的挑战:1)大规模、按帧标注的要求;2)两阶段方法的复杂性。为了解决这些挑战,我们引入了一种新的半监督框架,通过学习仅包含已标记图像数据集和未标记视频序列的实例跟踪网络。通过实例对比目标,我们学习了一种将实例与其他实例区分开来的嵌入方法。结果表明,即使只用图像训练,所学习的特征表示对于实例外观变化也是稳健的,因此能够跨帧稳定地跟踪目标。通过自监督的方式从未标记视频中学习对应关系,进一步增强了嵌入的跟踪能力。此外,我们将该模块集成到单级实例分割和姿态估计框架中,与两级网络相比,显著降低了跟踪的计算复杂度。我们在YouTube-VIS和PoseTrack数据集上进行了实验。在不需要任何视频标注的情况下,我们提出的方法可以获得与大多数全监督方法相当甚至更好的性能。
Tracking segmentation masks of multiple instances has been intensively studied, but still faces two fundamental challenges: 1) the requirement of large-scale, frame-wise annotation, and 2) the complexity of two-stage approaches. To resolve these challenges, we introduce a novel semisupervised framework by learning instance tracking networks with only a labeled image dataset and unlabeled video sequences. With an instance contrastive objective, we learn an embedding to discriminate each instance from the others. We show that even when only trained with images, the learned feature representation is robust to instance appearance variations, and is thus able to track objects steadily across frames. We further enhance the tracking capability of the embedding by learning correspondence from unlabeled videos in a self-supervised manner. In addition, we integrate this module into single-stage instance segmentation and pose estimation frameworks, which significantly reduce the computational complexity of tracking compared to two-stage networks. We conduct experiments on the YouTube-VIS and PoseTrack datasets. Without any video annotation efforts, our proposed method can achieve comparable or even better performance than most fullysupervised methods1.