Re-identification for Online Person Tracking by Modeling Space-Time Continuum

Re-identification for Online Person Tracking by Modeling Space-Time Continuum
复制标题

DOI:
10.1109/cvprw.2018.00193
复制
发表时间:
2018-06
期刊:
2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子:
--
通讯作者:
N. Narayan;Nishant Sankaran;S. Setlur;V. Govindaraju
N. Narayan;Nishant Sankaran;S. Setlur;V. Govindaraju
中科院分区:
其他
文献类型:
--
作者:
N. Narayan;Nishant Sankaran;S. Setlur;V. Govindaraju

文献摘要

被引文献

相似文献

提出了一种基于摄像机网络时空连续性学习的多人多摄像机跟踪新方法。在真实的场景中跟踪多个人所涉及的一些挑战包括a)确保所有人的可靠连续关联,以及B)考虑盲点或进入/退出点的存在。大多数现有的方法都设计了复杂的模型,需要大量的参数调整,这对于深度学习方法来说是一项重要的任务,因为它们不能直接应用于解决上述挑战。在这里,我们通过提出一种基于使用LSTM网络进行人员重新识别的判别式时空学习跟踪方法,以连贯的方式处理上述问题。当不知道关于个体的方面或个体的数量的先验信息时,这种方法更鲁棒。其想法是通过连续关联和从过去将不同个体与特定轨迹关联的错误中恢复来将检测识别为属于同一个体。我们利用LSTM注入时间信息的能力,通过联合整合视觉外观特征和位置信息来预测新检测属于同一跟踪实体的可能性。与CamNeT数据集上的先前最先进的方法相比,所提出的方法在错误率方面提高了50%,与DukeMTMC数据集上的基线方法相比,提高了18%。
We present a novel approach to multi-person multi-camera tracking based on learning the space-time continuum of a camera network. Some challenges involved in tracking multiple people in real scenarios include a) ensuring reliable continuous association of all persons, and b) accounting for presence of blind-spots or entry/exit points. Most of the existing methods design sophisticated models that require heavy tuning of parameters and it is a nontrivial task for deep learning approaches as they cannot be applied directly to address the above challenges. Here, we deal with the above points in a coherent way by proposing a discriminative spatio-temporal learning approach for tracking based on person re-identification using LSTM networks. This approach is more robust when no a-priori information about the aspect of an individual or the number of individuals is known. The idea is to identify detections as belonging to the same individual by continuous association and recovering from past errors in associating different individuals to a particular trajectory. We exploit LSTM's ability to infuse temporal information to predict the likelihood that new detections belong to the same tracked entity by jointly incorporating visual appearance features and location information. The proposed approach gives a 50% improvement in the error rate compared to the previous state-of-the-art method on the CamNeT dataset and 18% improvement as compared to the baseline approach on DukeMTMC dataset.