Learning a robust representation via a deep network on symmetric positive definite manifolds

Learning a robust representation via a deep network on symmetric positive definite manifolds
复制标题

通过对称正定流形上的深度网络学习鲁棒表示

DOI:
10.1016/j.patcog.2019.03.007
复制
发表时间:
2019-08-01
影响因子:
8
通讯作者:
Jia, Yunde
Jia, Yunde
中科院分区:
计算机科学1区
文献类型:
--
作者:
Gao, Zhi;Wu, Yuwei;Jia, Yunde

文献摘要

被引文献

相似文献

最近的研究表明,卷积神经网络(CNN)的卷积特征聚合可以在各种计算机视觉任务中获得令人印象深刻的性能。对称正定矩阵(SPD)由于其学习适当的统计表示来表征视觉特征的底层结构的卓越能力而成为一个强大的工具。在本文中,我们提出了一种在端到端深度网络下通过SPD生成和SPD变换将深度卷积特征聚合成鲁棒表示的方法。为此,在我们的方法中引入了几个新的层,包括非线性核生成层、矩阵变换层和向量变换层。采用非线性核生成层将卷积特征聚合成核矩阵,保证核矩阵为SPD矩阵。矩阵变换层的目的是将原SPD表示投影到一个更紧凑和判别的SPD流形。在矢量变换层进行向量化和归一化操作,取SPD表示的上三角元素,进行功率归一化和l(2)归一化,减少冗余,加快收敛速度。我们的网络中的SPD矩阵可以被认为是连接卷积特征和高级语义特征的中级表示。大量实验的结果表明,我们的方法明显优于最先进的方法。(C) 2019 Elsevier Ltd.版权所有。
Recent studies have shown that aggregating convolutional features of a Convolutional Neural Network (CNN) can obtain impressive performance for a variety of computer vision tasks. The Symmetric Positive Definite (SPD) matrix becomes a powerful tool due to its remarkable ability to learn an appropriate statistic representation to characterize the underlying structure of visual features. In this paper, we propose a method of aggregating deep convolutional features into a robust representation through the SPD generation and the SPD transformation under an end-to-end deep network. To this end, several new layers are introduced in our method, including a nonlinear kernel generation layer, a matrix transformation layer, and a vector transformation layer. The nonlinear kernel generation layer is employed to aggregate convolutional features into a kernel matrix which is guaranteed to be an SPD matrix. The matrix transformation layer is designed to project the original SPD representation to a more compact and discriminative SPD manifold. The vectorization and normalization operations are performed in the vector transformation layer to take the upper triangle elements of the SPD representation and carry out the power normalization and l(2) normalization to reduce the redundancy and accelerate the convergence. The SPD matrix in our network can be considered as a mid-level representation bridging convolutional features and high-level semantic features. Results of extensive experiments show that our method notably outperforms state-of-the-art methods. (C) 2019 Elsevier Ltd. All rights reserved.