Image to Video Person Re-Identification by Learning Heterogeneous Dictionary Pair With Feature Projection Matrix

Image to Video Person Re-Identification by Learning Heterogeneous Dictionary Pair With Feature Projection Matrix
复制标题

通过学习异构字典对和特征投影矩阵进行图像到视频的人物重新识别

DOI:
10.1109/tifs.2017.2765524
复制
发表时间:
2018-03-01
影响因子:
6.8
通讯作者:
Zheng, Wei-Shi
Zheng, Wei-Shi
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhu, Xiaoke;Jing, Xiao-Yuan;Zheng, Wei-Shi

文献摘要

被引文献

相似文献

在视频监控和取证应用中,人员身份识别起着重要的作用。在许多情况下,需要在图像和视频片段之间进行人物重新识别,例如,从大量行人视频中重新识别嫌疑人,给出嫌疑人的单个图像。在这种情况下,我们将重新识别称为图像到视频的人重新识别(IVPR)。在实际应用中,图像和视频通常用不同的特征来表示,并且每个视频内的帧之间通常存在较大的变化。这些因素使得图像和视频之间的匹配成为一个非常具有挑战性的任务。本文提出了一种联合<inline-formula><tex-math notation="LaTeX">特征投影</tex-math></inline-formula>矩阵和异<inline-formula><tex-math notation="LaTeX">时特征</tex-math></inline-formula>对<inline-formula><tex-math notation="LaTeX">学习</tex-math></inline-formula>算法。<inline-formula><tex-math notation="LaTeX"></tex-math></inline-formula>具体而言,PHDL联合学习视频内投影矩阵和一对异构图像和视频字典。利用学习的投影矩阵,可以减少每个视频内的变化对匹配造成的影响。利用学习的字典对,可以将异构图像和视频特征转换为具有相同维数的编码系数,从而可以利用编码系数进行匹配。此外,为了保证获得的编码系数具有良好的鉴别能力,PPDL设计了一个点到集的系数鉴别项。为了更好地利用行人视频数据中包含的互补时空和视觉外观信息,我们进一步提出了一种多视图PPDL方法,该方法可以在字典学习过程中有效地融合不同的视频信息。在四个公开的人物序列数据集上的实验证明了所提方法的有效性。
Person re-identification plays an important role in video surveillance and forensics applications. In many cases, person re-identification needs to be conducted between image and video clip, e.g., re-identifying a suspect from large quantities of pedestrian videos given a single image of the suspect. We call re-identification in this scenario as image to video person re-identification (IVPR). In practice, image and video are usually represented with different features, and there usually exist large variations between frames within each video. These factors make matching between image and video become a very challenging task. In this paper, we propose a joint feature <inline-formula> <tex-math notation="LaTeX">${p}$ </tex-math></inline-formula>rojection matrix and <inline-formula> <tex-math notation="LaTeX">${h}$ </tex-math></inline-formula>eterogeneous <inline-formula> <tex-math notation="LaTeX">${d}$ </tex-math></inline-formula>ictionary pair <inline-formula> <tex-math notation="LaTeX">${l}$ </tex-math></inline-formula>earning (PHDL) approach for IVPR. Specifically, the PHDL jointly learns an intra-video projection matrix and a pair of heterogeneous image and video dictionaries. With the learned projection matrix, the influence caused by the variations within each video on the matching can be reduced. With the learned dictionary pair, the heterogeneous image and video features can be transformed into coding coefficients with the same dimension, such that the matching can be conducted by using the coding coefficients. Furthermore, to ensure that the obtained coding coefficients own favorable discriminability, the PHDL designs a point-to-set coefficient discriminant term. To make better use of the complementary spatial-temporal and visual appearance information contained in pedestrian video data, we further propose a multi-view PHDL approach, which can fuse different video information effectively in the dictionary learning process. Experiments on four publicly available person sequence data sets demonstrate the effectiveness of the proposed approaches.