Deep feature learning with relative distance comparison for person re-identification

Deep feature learning with relative distance comparison for person re-identification
复制标题

通过相对距离比较进行深度特征学习以进行行人重新识别

DOI:
10.1016/j.patcog.2015.04.005
复制
发表时间:
2015-10-01
影响因子:
8
通讯作者:
Chao, Hongyang
Chao, Hongyang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ding, Shengyong;Lin, Liang;Chao, Hongyang

文献摘要

被引文献

相似文献

在智能视频监控中,识别不同场景中的同一个人是一项重要而困难的任务。它的主要困难在于如何在区分不同个体的同时,保持同一个人的相似性,以对抗大的外观和结构变化。在本文中,我们提出了一个基于深度神经网络的可扩展距离驱动特征学习框架,用于人员重新识别,并证明了其处理现有挑战的有效性。具体来说,给定带有类标签(人物ID)的训练图像,我们首先产生大量的三元组单元,每个单元包含三个图像,即一个人具有匹配的参考和一个不匹配的参考。将单元作为输入,我们构建卷积神经网络来生成分层表示,然后使用L2距离度量。通过参数优化,我们的框架倾向于最大化每个三重态单元的匹配对和错配对之间的相对距离。此外,一个不平凡的问题与框架是,三元组组织立方地扩大训练三元组的数量,因为一个图像可以涉及到几个三元组单元。为了克服这个问题,我们开发了一个有效的三元组生成方案和优化的梯度下降算法,使计算量主要取决于原始图像的数量,而不是三元组的数量。在几个具有挑战性的数据库上,我们的方法取得了非常有希望的结果,并优于其他最先进的方法。(C)2015爱思唯尔有限公司版权所有。
Identifying the same individual across different scenes is an important yet difficult task in intelligent video surveillance. Its main difficulty lies in how to preserve similarity of the same person against large appearance and structure variation while discriminating different individuals. In this paper, we present a scalable distance driven feature learning framework based on the deep neural network for person re-identification, and demonstrate its effectiveness to handle the existing challenges. Specifically, given the training images with the class labels (person IDs), we first produce a large number of triplet units, each of which contains three images, i.e. one person with a matched reference and a mismatched reference. Treating the units as the input, we build the convolutional neural network to generate the layered representations, and follow with the L2 distance metric. By means of parameter optimization, our framework tends to maximize the relative distance between the matched pair and the mismatched pair for each triplet unit. Moreover, a nontrivial issue arising with the framework is that the triplet organization cubically enlarges the number of training triplets, as one image can be involved into several triplet units. To overcome this problem, we develop an effective triplet generation scheme and an optimized gradient descent algorithm, making the computational load mainly depend on the number of original images instead of the number of triplets. On several challenging databases, our approach achieves very promising results and outperforms other state-of-the-art approaches. (C) 2015 Elsevier Ltd. All rights reserved.