Attention adjacency matrix based graph convolutional networks for skeleton-based action recognition

Attention adjacency matrix based graph convolutional networks for skeleton-based action recognition
复制标题

基于注意力邻接矩阵的图卷积网络用于基于骨架的动作识别

DOI:
10.1016/j.neucom.2021.02.001
复制
发表时间:
2021
期刊:
影响因子:
6
通讯作者:
Xuesong Gao
Xuesong Gao
中科院分区:
计算机科学2区
文献类型:
--
作者:
Jun Xie;Qiguang Miao;Ruyi Liu;Wentian Xin;Lei Tang;Sheng Zhong;Xuesong Gao

文献摘要

相似文献

在图形卷积网络(GCN)的推动下,人类行为识别方面的最新进展是实质性的。然而,图卷积核的设计策略带来了两个主要问题:一是图顶点的邻点集划分策略依赖于人工设计的重心,在动作识别中对不同骨架的泛化能力有限;二是现有的基于GCN的方法只能捕获关节之间的局部物理依赖关系,并且由于过度平滑而导致丢失隐含的关节关联。在这项工作中,我们提出(1)一个新的关注邻接矩阵(AAM)来设计图的卷积核和(2)一个维度-关注块来提高模型的稳健性。该算法采用一种新的邻域划分策略,将邻接矩阵分解为若干个参数矩阵。同时,在生成注意矩阵的过程中引入了注意机制。通过ResNet将矩阵和参数矩阵组合成AAM,进一步提出了基于AAM的图卷积网络(AAM-GCN)。提出的维度注意块通过扩展通道注意的思想,强化了骨架数据各维度中的重要信息。在NTU-RGB+D和Kinetics两个大规模数据集上的大量实验表明,AAM-GCN的性能优于目前最先进的工作。
Recent progress on human action recognition, fueled by the Graph Convolutional Network (GCN), has been substantial. However, two main problems are caused by the design strategy of graph convolution kernels: first, the partitioning strategy of neighbor set for graph vertices relies on the gravity center designed manually, which is limited in generalizability to diverse skeletons in action recognition; second, the existing GCN-based methods can only capture local physical dependencies among joints and result in missing implicit joint correlations due to over-smoothing. In this work, we present (1) a novel attention adjacency matrix (AAM) to design graph convolution kernels and (2) a dimension-attention block to improve the robustness of the model. Specifically, the proposed AAM is designed by a novel partitioning strategy for the neighbor set, through which an adjacency matrix is decomposed into several parametric matrices. Simultaneously, attention mechanism is introduced in the process to generate an attention matrix. Combining the matrix and the parametric matrices into an AAM through ResNet, we further exhibit the AAM based graph convolution network (AAM-GCN). The proposed dimension-attention block strengthens the important information in each dimension of skeleton data by extending the idea of channel-attention. Extensive experiments on two large-scale datasets, NTU-RGB+D and Kinetics, demonstrate that AAM-GCN achieves better performance than the state-of-the-art works.