A Two-Stage Minimum Cost Multicut Approach to Self-supervised Multiple Person Tracking

A Two-Stage Minimum Cost Multicut Approach to Self-supervised Multiple Person Tracking
复制标题

DOI:
10.1007/978-3-030-69532-3_33
复制
发表时间:
2020
期刊:
2018 4th International Conference on Universal Village (UV)
影响因子:
--
通讯作者:
Kalun Ho;Amirhossein Kardoost;F. Pfreundt;J. Keuper;M. Keuper
Kalun Ho;Amirhossein Kardoost;F. Pfreundt;J. Keuper;M. Keuper
中科院分区:
其他
文献类型:
--
作者:
Kalun Ho;Amirhossein Kardoost;F. Pfreundt;J. Keuper;M. Keuper

文献摘要

相似文献

多目标跟踪(MOT)是计算机视觉领域的一项长期任务。基于检测跟踪范例的当前方法需要某种领域知识或监督来将数据正确地关联到轨道中。在这项工作中,我们提出了一个自我监督的多目标跟踪方法的基础上的视觉特征和最小成本提升多割。我们的方法是基于直接的时空线索,可以从相邻帧的图像序列中提取,而无需监督。基于这些线索的聚类使我们能够学习手头跟踪任务所需的外观不变性,并训练AutoEncoder生成合适的潜在表示。因此,所得到的潜在表示可以作为鲁棒的外观线索,用于跟踪,即使在大的时间距离,其中没有可靠的时空特征可以被提取。我们表明,尽管在没有使用所提供的注释的情况下进行了训练,但我们的模型在具有挑战性的MOT Benchmark行人跟踪上提供了有竞争力的结果。
Multiple Object Tracking (MOT) is a long-standing task in computer vision. Current approaches based on the tracking by detection paradigm either require some sort of domain knowledge or supervision to associate data correctly into tracks. In this work, we present a self-supervised multiple object tracking approach based on visual features and minimum cost lifted multicuts. Our method is based on straight-forward spatio-temporal cues that can be extracted from neighboring frames in an image sequences without supervision. Clustering based on these cues enables us to learn the required appearance invariances for the tracking task at hand and train an AutoEncoder to generate suitable latent representations. Thus, the resulting latent representations can serve as robust appearance cues for tracking even over large temporal distances where no reliable spatio-temporal features can be extracted. We show that, despite being trained without using the provided annotations, our model provides competitive results on the challenging MOT Benchmark for pedestrian tracking.