Face recognition on large-scale video in the wild with hybrid Euclidean-and-Riemannian metric learning

Face recognition on large-scale video in the wild with hybrid Euclidean-and-Riemannian metric learning
复制标题

DOI:
10.1016/j.patcog.2015.03.011
复制
发表时间:
2015-10
期刊:
Pattern Recognit.
影响因子:
--
通讯作者:
Zhiwu Huang;Ruiping Wang;S. Shan;Xilin Chen
Zhiwu Huang;Ruiping Wang;S. Shan;Xilin Chen
中科院分区:
其他
文献类型:
--
作者:
Zhiwu Huang;Ruiping Wang;S. Shan;Xilin Chen

文献摘要

被引文献

相似文献

由于监控摄像头、手持设备、互联网上传和其他来源捕获的视频数据无处不在,野外大规模视频的人脸识别变得越来越重要。通过将每个视频视为一个图像集,基于集合的方法最近在基于视频的人脸识别领域取得了巨大的成功。在现实世界中,视频往往包含极其复杂的数据变化,因此对基于集合的方法的集合建模提出了很大的挑战。本文提出了一种新的混合欧氏-黎曼度量学习(HERML)方法来融合图像集的多种统计信息。具体地说,我们同时表示每个图像集的均值,协方差矩阵和高斯分布,这通常是相辅相成的集合建模方面。然而,这是不平凡的融合,因为平均值,协方差矩阵和高斯模型通常位于多个异构空间配备了欧几里得或黎曼度量。因此,我们首先隐式映射到高维希尔伯特空间的原始统计利用欧氏和黎曼核。通过基于LogDet散度的目标函数,混合核然后通过我们的混合度量学习框架进行融合,该框架可以有效地对大规模视频进行融合过程。该方法在四个公共和具有挑战性的大规模视频人脸数据集上进行了评估。大量的实验结果表明,我们的方法有一个明显的优越性,比国家的最先进的基于集合的方法为基础的大规模视频人脸识别。
Face recognition on large-scale video in the wild is becoming increasingly important due to the ubiquity of video data captured by surveillance cameras, handheld devices, Internet uploads, and other sources. By treating each video as one image set, set-based methods recently have made great success in the field of video-based face recognition. In the wild world, videos often contain extremely complex data variations and thus pose a big challenge of set modeling for set-based methods. In this paper, we propose a novel Hybrid Euclidean-and-Riemannian Metric Learning (HERML) method to fuse multiple statistics of image set. Specifically, we represent each image set simultaneously by mean, covariance matrix and Gaussian distribution, which generally complement each other in the aspect of set modeling. However, it is not trivial to fuse them since mean, covariance matrix and Gaussian model typically lie in multiple heterogeneous spaces equipped with Euclidean or Riemannian metric. Therefore, we first implicitly map the original statistics into high dimensional Hilbert spaces by exploiting Euclidean and Riemannian kernels. With a LogDet divergence based objective function, the hybrid kernels are then fused by our hybrid metric learning framework, which can efficiently perform the fusing procedure on large-scale videos. The proposed method is evaluated on four public and challenging large-scale video face datasets. Extensive experimental results demonstrate that our method has a clear superiority over the state-of-the-art set-based methods for large-scale video-based face recognition.