Temporal-stochastic tensor features for action recognition

Temporal-stochastic tensor features for action recognition
复制标题

DOI:
10.1016/j.mlwa.2022.100407
复制
发表时间:
2022-09
影响因子:
--
通讯作者:
Bojan Batalo;L. S. Souza;B. Gatto;Naoya Sogi;K. Fukui
Bojan Batalo;L. S. Souza;B. Gatto;Naoya Sogi;K. Fukui
中科院分区:
--
文献类型:
--
作者:
Bojan Batalo;L. S. Souza;B. Gatto;Naoya Sogi;K. Fukui

文献摘要

相似文献

在本文中,我们提出了时间随机产品格拉斯曼流形(TS-PGM),一种有效的方法,张量分类的任务,如手势和动作识别。我们的方法建立在将张量表示为乘积格拉斯曼流形(PGM)上的点的思想上。这是通过将张量模式映射到线性子空间来实现的,其中每个子空间可以被视为对应模式的格拉斯曼流形(GM)上的一个点。随后,可以通过PGM以自然的方式统一各个模式的因子流形。然而,这种方法可能会通过平等对待所有模式而丢弃有区别的信息,并且不考虑诸如视频的时间张量的性质。因此,我们引入时间随机张量特征(TST特征)来从张量中提取时间信息,并将其编码在序列保持的TST子空间中。这些特征和常规张量模式可以同时用于PGM。我们的框架解决了时间张量的分类问题,同时继承了PGM的统一数学解释,因为TST子空间可以自然地集成到PGM作为一个新的因子流形。此外,我们从两个方面增强了我们的方法:(1)我们通过将子空间投影到广义差子空间上来提高区分能力,(2)我们利用核映射来构造能够处理非线性数据分布的核化子空间。手势和动作识别数据集上的实验结果表明,我们的方法基于子空间表示与显式TST功能优于纯粹的时空方法。
In this paper, we propose Temporal-Stochastic Product Grassmann Manifold (TS-PGM), an efficient method for tensor classification in tasks such as gesture and action recognition. Our approach builds on the idea of representing tensors as points on Product Grassmann Manifold (PGM). This is achieved by mapping tensor modes to linear subspaces, where each subspace can be seen as a point on a Grassmann Manifold (GM) of the corresponding mode. Subsequently, it is possible to unify factor manifolds of respective modes in a natural way via PGM. However, this approach possibly discards discriminative information by treating all modes equally, and not considering the nature of temporal tensors such as videos. Therefore, we introduce Temporal-Stochastic Tensor features (TST features) to extract temporal information from tensors and encode them in a sequence-preserving TST subspace. These features and regular tensor modes can then be simultaneously used on PGM. Our framework addresses the problem of classification of temporal tensors while inheriting the unified mathematical interpretation of PGM because the TST subspace can be naturally integrated into PGM as a new factor manifold. Additionally, we enhance our method in two ways: (1) we improve the discrimination ability by projecting subspaces onto a Generalized Difference Subspace, and (2) we utilize kernel mapping to construct kernelized subspaces able to handle nonlinear data distribution. Experimental results on gesture and action recognition datasets show that our methods based on subspace representation with explicit TST features outperform pure spatio-temporal approaches.