Image to Video Person Re-Identification by Learning Heterogeneous Dictionary Pair With Feature Projection Matrix
Image to Video Person Re-Identification by Learning Heterogeneous Dictionary Pair With Feature Projection Matrix
复制标题
通过学习异构字典对和特征投影矩阵进行图像到视频的人物重新识别
DOI:
10.1109/tifs.2017.2765524
复制
发表时间:
2018-03-01
影响因子:
6.8
通讯作者:
Zheng, Wei-Shi
中科院分区:
文献类型:
--
作者:
Zhu, Xiaoke;Jing, Xiao-Yuan;Zheng, Wei-Shi
Person re-identification plays an important role in video surveillance and forensics applications. In many cases, person re-identification needs to be conducted between image and video clip, e.g., re-identifying a suspect from large quantities of pedestrian videos given a single image of the suspect. We call re-identification in this scenario as image to video person re-identification (IVPR). In practice, image and video are usually represented with different features, and there usually exist large variations between frames within each video. These factors make matching between image and video become a very challenging task. In this paper, we propose a joint feature <inline-formula> <tex-math notation="LaTeX">${p}$ </tex-math></inline-formula>rojection matrix and <inline-formula> <tex-math notation="LaTeX">${h}$ </tex-math></inline-formula>eterogeneous <inline-formula> <tex-math notation="LaTeX">${d}$ </tex-math></inline-formula>ictionary pair <inline-formula> <tex-math notation="LaTeX">${l}$ </tex-math></inline-formula>earning (PHDL) approach for IVPR. Specifically, the PHDL jointly learns an intra-video projection matrix and a pair of heterogeneous image and video dictionaries. With the learned projection matrix, the influence caused by the variations within each video on the matching can be reduced. With the learned dictionary pair, the heterogeneous image and video features can be transformed into coding coefficients with the same dimension, such that the matching can be conducted by using the coding coefficients. Furthermore, to ensure that the obtained coding coefficients own favorable discriminability, the PHDL designs a point-to-set coefficient discriminant term. To make better use of the complementary spatial-temporal and visual appearance information contained in pedestrian video data, we further propose a multi-view PHDL approach, which can fuse different video information effectively in the dictionary learning process. Experiments on four publicly available person sequence data sets demonstrate the effectiveness of the proposed approaches.