Grassmannian learning mutual subspace method for image set recognition

Grassmannian learning mutual subspace method for image set recognition
复制标题

DOI:
10.1016/j.neucom.2022.10.040
复制
发表时间:
2021-11
期刊:
影响因子:
6
通讯作者:
L. S. Souza;Naoya Sogi;B. Gatto;Takumi Kobayashi;K. Fukui
L. S. Souza;Naoya Sogi;B. Gatto;Takumi Kobayashi;K. Fukui
中科院分区:
计算机科学2区
文献类型:
--
作者:
L. S. Souza;Naoya Sogi;B. Gatto;Takumi Kobayashi;K. Fukui

文献摘要

相似文献

本文讨论了给定一组图像作为输入的对象识别问题(例如,多个摄像机源和视频帧)。基于卷积神经网络(CNN)的框架不能有效地利用这些集合,处理观察到的模式,不捕获底层特征分布,因为它不考虑集合中图像的方差。为了解决这个问题,我们提出了格拉斯曼学习互子空间方法(G-LMSM),这是一个嵌入在CNN之上的NN层,可以更有效地处理图像集,并且可以以端到端的方式进行训练。首先用一个低维的输入子空间表示图像集,然后用字典子空间与输入子空间的正则角的相似性进行匹配,这是一个可解释且易于计算的度量。G-LMSM的关键思想是字典子空间被学习为Grassmann流形上的点,并使用黎曼随机梯度下降进行优化。这种学习是稳定的,有效的,理论上是有根据的。我们证明了我们所提出的方法的有效性,手形识别,人脸识别,面部表情识别。
This paper addresses the problem of object recognition given a set of images as input (e.g., multiple camera sources and video frames). Convolutional neural network (CNN)-based frameworks do not exploit these sets effectively, processing a pattern as observed, not capturing the underlying feature distribution as it does not consider the variance of images in the set. To address this issue, we propose the Grassmannian learning mutual subspace method (G-LMSM), a NN layer embedded on top of CNNs that can process image sets more effectively and can be trained in an end-to-end manner. The image set is first represented by a low-dimensional input subspace and then this input subspace is matched with dictionary subspaces by a similarity of their canonical angles, an interpretable and easy to compute metric. The key idea of G-LMSM is that the dictionary subspaces are learned as points on the Grassmann manifold, optimized with Riemannian stochastic gradient descent. This learning is stable, efficient and theoretically well-grounded. We demonstrate the effectiveness of our proposed method on hand shape recognition, face identification, and facial emotion recognition.