Image-to-Video Person Re-Identification by Reusing Cross-modal Embeddings

Image-to-Video Person Re-Identification by Reusing Cross-modal Embeddings
复制标题

DOI:
--
复制
发表时间:
2018-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhongwei Xie;Lin Li;X. Zhong;Luo Zhong
Zhongwei Xie;Lin Li;X. Zhong;Luo Zhong
中科院分区:
其他
文献类型:
--
作者:
Zhongwei Xie;Lin Li;X. Zhong;Luo Zhong

文献摘要

被引文献

相似文献

图像到视频人重新识别通过来自由非重叠相机捕获的大量行人视频的探测图像来识别目标人。尽管取得了很大的进展,它仍然是具有挑战性的多模态场景,即图像和视频之间的匹配。目前,最先进的方法主要集中在特定的任务数据,忽略了不同的,但相关的任务的额外信息。在本文中,我们提出了一个端到端的神经网络框架,通过利用从额外信息中学习的跨模态嵌入来进行图像到视频的人物重新识别。具体来说,来自图像字幕和视频字幕模型的跨模态嵌入被重用,以帮助学习的特征被投影到一个协调的空间中,在那里可以直接计算相似度。此外,固定模型重用方法的训练步骤被集成到我们的框架中,它可以包含有益的信息,并最终使目标网络独立于现有的模型。除此之外,我们提出的框架采用CNN和LSTM来提取视觉和时空特征,并结合识别和验证模型的优势来提高学习特征的区分能力。实验结果表明,我们的框架缩小了异构数据之间的差距,并获得明显的改善,在图像到视频的人重新识别的有效性。
Image-to-video person re-identification identifies a target person by a probe image from quantities of pedestrian videos captured by non-overlapping cameras. Despite the great progress achieved,it's still challenging to match in the multimodal scenario,i.e. between image and video. Currently,state-of-the-art approaches mainly focus on the task-specific data,neglecting the extra information on the different but related tasks. In this paper,we propose an end-to-end neural network framework for image-to-video person reidentification by leveraging cross-modal embeddings learned from extra information.Concretely speaking,cross-modal embeddings from image captioning and video captioning models are reused to help learned features be projected into a coordinated space,where similarity can be directly computed. Besides,training steps from fixed model reuse approach are integrated into our framework,which can incorporate beneficial information and eventually make the target networks independent of existing models. Apart from that,our proposed framework resorts to CNNs and LSTMs for extracting visual and spatiotemporal features,and combines the strengths of identification and verification model to improve the discriminative ability of the learned feature. The experimental results demonstrate the effectiveness of our framework on narrowing down the gap between heterogeneous data and obtaining observable improvement in image-to-video person re-identification.