Learning Log-Determinant Divergences for Positive Definite Matrices

Learning Log-Determinant Divergences for Positive Definite Matrices
复制标题

DOI:
10.1109/tpami.2021.3073588
复制
发表时间:
2021-04
影响因子:
23.6
通讯作者:
A. Cherian;P. Stanitsas;Jue Wang;Mehrtash Harandi;V. Morellas;N. Papanikolopoulos
A. Cherian;P. Stanitsas;Jue Wang;Mehrtash Harandi;V. Morellas;N. Papanikolopoulos
中科院分区:
计算机科学1区
文献类型:
--
作者:
A. Cherian;P. Stanitsas;Jue Wang;Mehrtash Harandi;V. Morellas;N. Papanikolopoulos

文献摘要

相似文献

对称正定(SPD)矩阵形式的表示已经在各种视觉学习应用中得到推广,因为它们具有捕获视觉数据的丰富二阶统计的能力。有几个相似性的措施比较SPD矩阵与记录的好处。然而,为特定问题选择适当的衡量标准仍然是一个挑战,在大多数情况下,是试错过程的结果。在本文中,我们建议以数据驱动的方式学习相似性度量。为此,我们利用了$\alpha \beta$αβ-log-det发散,这是一个由标量$\alpha$α和$\beta$β参数化的元发散,包含了一个广泛的家庭流行的信息发散的SPD矩阵的这些参数的不同和离散的值。我们的关键思想是将这些参数放在一个连续体中,并从数据中学习它们。我们系统地扩展了这个想法,学习向量值参数,从而增加了底层非线性度量的表达能力。我们将发散学习问题与机器学习中的几个标准任务结合起来,包括有监督的判别字典学习和无监督的SPD矩阵聚类。我们提出了黎曼梯度下降方案,有效地优化我们的配方,并显示我们的方法对八个标准的计算机视觉任务的实用性。
Representations in the form of Symmetric Positive Definite (SPD) matrices have been popularized in a variety of visual learning applications due to their demonstrated ability to capture rich second-order statistics of visual data. There exist several similarity measures for comparing SPD matrices with documented benefits. However, selecting an appropriate measure for a given problem remains a challenge and in most cases, is the result of a trial-and-error process. In this paper, we propose to learn similarity measures in a data-driven manner. To this end, we capitalize on the $\alpha \beta$αβ-log-det divergence, which is a meta-divergence parametrized by scalars $\alpha$α and $\beta$β, subsuming a wide family of popular information divergences on SPD matrices for distinct and discrete values of these parameters. Our key idea is to cast these parameters in a continuum and learn them from data. We systematically extend this idea to learn vector-valued parameters, thereby increasing the expressiveness of the underlying non-linear measure. We conjoin the divergence learning problem with several standard tasks in machine learning, including supervised discriminative dictionary learning and unsupervised SPD matrix clustering. We present Riemannian gradient descent schemes for optimizing our formulations efficiently, and show the usefulness of our method on eight standard computer vision tasks.