Robust Video-Based Person Re-Identification by Hierarchical Mining

Robust Video-Based Person Re-Identification by Hierarchical Mining
复制标题

DOI:
10.1109/tcsvt.2021.3076097
复制
发表时间:
2021-04
影响因子:
8.4
通讯作者:
Zhikang Wang;Lihuo He;X. Tu;Jian Zhao;Xinbo Gao;Shengmei Shen;Jiashi Feng
Zhikang Wang;Lihuo He;X. Tu;Jian Zhao;Xinbo Gao;Shengmei Shen;Jiashi Feng
中科院分区:
工程技术1区
文献类型:
--
作者:
Zhikang Wang;Lihuo He;X. Tu;Jian Zhao;Xinbo Gao;Shengmei Shen;Jiashi Feng

文献摘要

相似文献

基于视频的人重新识别(Re-ID)旨在通过非重叠摄像机的视频序列来检索人。由于视点、姿势和遮挡随时间的变化,行人的某些特征在帧内不是连续的。然而,现有方法忽略了这种数据特性,并且网络倾向于仅学习视频序列中的帧之间的那些显著的连续特征。因此,学习的表征不能涵盖行人的所有特征,从而缺乏完整性和区分性。针对这一问题,我们提出了一种新的深层结构--分层挖掘网络(HMN),它通过参考时间和类内知识来挖掘尽可能多的行人特征。它由一个新的注意时间模块(ATM)和一个动态监督分支(DSB)组成,并利用平衡三重态损失(BTL)辅助训练。所提出的ATM具有行人感知能力,能够通过时间分析来评估每个特征的激活情况,从而更好地聚合行人的时间分散特征,从而消除受污染的特征。然后,DSB和BTL一起通过多重监督进一步增强了陈述的完整性。具体地说,DSB感知每个小批次中类内样本的差异,并为它们生成有针对性的监督信号,在这个过程中,BTL保证信号具有较小的类内变化和较大的类间变化。在两个基于视频的数据集MARS和DukeMTMC-VideoReID上的综合实验证明了每个组件的贡献以及所提出的HMN相对于最先进的HMN的优越性。在Market1501、DukeMTMC-Reid和MSMT17这三个流行的基于图像的数据集上对我们的模型进行了基准测试,进一步验证了所提出的DSB和BTL具有良好的泛化能力。
Video-based person re-identification (Re-ID) aims at retrieving the person through the video sequences across non-overlapping cameras. Some characteristics of pedestrians are not consecutive across frames due to the variations of viewpoints, postures, and occlusions over time. However, existing methods ignore such data peculiarity and the networks tend to only learn those salient consecutive characteristics among frames in video sequences. As a result, the learned representations fail to cover all the characteristics of pedestrians, thus lacking integrity and discrimination. To tackle this problem, we present a novel deep architecture termed Hierarchical Mining Network (HMN), which mines as many pedestrians’ characteristics by referring to the temporal and intra-class knowledge. It consists of a novel Attentive Temporal Module (ATM) and a Dynamic Supervising Branch (DSB), with a Balancing Triplet Loss (BTL) assisting the training. The proposed ATM, with pedestrian perceiving capacity, is capable of evaluating each activation of features through temporal analysis, so that the temporally scattered characteristics of pedestrians can be better aggregated and the contaminated ones can be eliminated. Then, the DSB along with the BTL further enhances the integrity of representations by multiple supervision. Specifically, the DSB perceives the diversities of intra-class samples in each mini-batch and generates targeted supervising signals for them, in which process the BTL guarantees the signals with smaller intra-class variations and larger inter-class variations. Comprehensive experiments on two video-based datasets, i.e., MARS, and DukeMTMC-VideoReID, demonstrate the contribution of each component and the superiority of the proposed HMN over the state-of-the-arts. Benchmarking our model on three popular image-based datasets, i.e., Market1501, DukeMTMC-Reid, and MSMT17 additionally verifies the promising generalizability of the proposed DSB and BTL.