Information disentanglement based cross-modal representation learning for visible-infrared person re-identification

Information disentanglement based cross-modal representation learning for visible-infrared person re-identification
复制标题

DOI:
10.1007/s11042-022-13669-3
复制
发表时间:
2022-09
影响因子:
3.6
通讯作者:
Xiaoke Zhu;Minghao Zheng;Xiaopan Chen;Xinyu Zhang;Caihong Yuan;Fan Zhang
Xiaoke Zhu;Minghao Zheng;Xiaopan Chen;Xinyu Zhang;Caihong Yuan;Fan Zhang
中科院分区:
计算机科学4区
文献类型:
--
作者:
Xiaoke Zhu;Minghao Zheng;Xiaopan Chen;Xinyu Zhang;Caihong Yuan;Fan Zhang

文献摘要

相似文献

可见光-红外身份识别(VI-ReID)是自动视频监控和取证中的一项重要而又极具挑战性的任务。虽然现有的VI-Reid方法已经取得了非常令人鼓舞的结果,但如何充分利用交叉通道可见光和红外图像中包含的有用信息还没有得到很好的研究。本文针对VI-REID提出了一种基于信息解缠的跨模式表示学习方法(IDCRL)。具体地,IDCRL首先分别使用共享特征学习模块和特定特征学习模块从每个通道的数据中提取共享特征和特定特征。为了确保共享的和特定的信息能够被很好地解开,我们对每个通道的共享和特定的特征施加了正交性约束。为了使从同一人的可见光和红外图像中提取的共享特征具有较高的相似性,IDCRL设计了共享特征一致性约束。此外,IDCRL使用通道感知损失来确保可以有效地从每个通道中提取有用的通道特定特征。然后,将所获得的共享特征和特定特征连接起来作为每个图像的表示。最后,利用同一性损失函数和跨模式判别损失函数来提高图像表示的可区分性。我们在基准的可见光-红外行人数据集(SYSU-MM01和RegDB)上进行了全面的实验,以评估IDCRL方法的有效性。实验结果表明,IDCRL方法的性能优于同类方法。在SYSU-MM01数据集上,我们的方法在全搜索和室内模式下的一阶匹配率分别达到了62.35%和71.64%。在RegDB数据集上,我们的方法在可见光到热模式和热到可见光模式下的排序结果分别达到了76.32%和75.49%。
Visible-infrared person re-identification (VI-ReID) is an important but very challenging task in the automated video surveillance and forensics. Although existing VI-ReID methods have achieved very encouraging results, how to make full use of the useful information contained in cross-modality visible and infrared images has not been well studied. In this paper, we propose an Information Disentanglement based Cross-modal Representation Learning (IDCRL) approach for VI-ReID. Specifically, IDCRL first extracts the shared and specific features from data of each modality by using the shared feature learning module and the specific feature learning module, respectively. To ensure that the shared and specific information can be well disentangled, we impose an orthogonality constraint on the shared and specific features of each modality. To make the shared features extracted from the visible and infrared images of the same person own high similarity, IDCRL designs a shared feature consistency constraint. Furthermore, IDCRL uses a modality-aware loss to ensure that the useful modality-specific features can be extracted from each modality effectively. Then, the obtained shared and specific features are concatenated as the representation of each image. Finally, identity loss function and cross-modal discriminant loss function are employed to enhance the discriminability of the obtained image representation. We conducted comprehensive experiments on the benchmark visible-infrared pedestrian datasets (SYSU-MM01 and RegDB) to evaluate the efficacy of our IDCRL approach. Experimental results demonstrate that IDCRL outperforms the compared state-of-the-art methods. On the SYSU-MM01 dataset, the rank-1 matching rate of our approach reaches 62.35% and 71.64% in the all-search and in-door modes, respectively. On the RegDB dataset, the rank-1 result of our approach reaches 76.32% and 75.49% in the visible to thermal and thermal to visible modes, respectively.