Part Aware Contrastive Learning for Self-Supervised Action Recognition

Part Aware Contrastive Learning for Self-Supervised Action Recognition
复制标题

DOI:
10.48550/arxiv.2305.00666
复制
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Yilei Hua;Wen-Lan Wu;Ce Zheng;Aidong Lu;Mengyuan Liu;Chen Chen-Chen;Shiqian Wu
Yilei Hua;Wen-Lan Wu;Ce Zheng;Aidong Lu;Mengyuan Liu;Chen Chen-Chen;Shiqian Wu
中科院分区:
其他
文献类型:
--
作者:
Yilei Hua;Wen-Lan Wu;Ce Zheng;Aidong Lu;Mengyuan Liu;Chen Chen-Chen;Shiqian Wu

文献摘要

相似文献

近年来,基于骨架序列和对比学习的自监督动作识别研究取得了显著的成果。已经观察到,人类动作特征的语义区别通常由身体局部部分来表示,例如腿或手,这有利于基于骨骼的动作识别。提出了一种基于注意力的骨架表征对比学习框架SkeAttnCLR,该框架综合了基于骨架的动作表征的局部相似性和全局特征。为此,使用多头注意掩模模块从骨架中学习软注意掩码特征,抑制不显著的局部特征,同时强调局部显著特征,从而使相似的局部特征在特征空间中更接近。此外,通过基于具有全局特征的显著和非显著特征扩展对比对来生成充足的对比对,从而引导网络学习整个骨架的语义表示。因此,通过注意掩蔽机制,SkeAttnCLR学习不同数据增强视图下的局部特征。实验结果表明,局部特征相似度的加入显著提高了基于骨架的动作表示。我们提出的SkeAttnCLR在NTURGB+D、NTU120-RGB+D和PKU-MMD数据集上的性能优于最先进的方法。代码和设置可在以下存储库获得:https://github.com/GitHubOfHyl97/SkeAttnCLR.
In recent years, remarkable results have been achieved in self-supervised action recognition using skeleton sequences with contrastive learning. It has been observed that the semantic distinction of human action features is often represented by local body parts, such as legs or hands, which are advantageous for skeleton-based action recognition. This paper proposes an attention-based contrastive learning framework for skeleton representation learning, called SkeAttnCLR, which integrates local similarity and global features for skeleton-based action representations. To achieve this, a multi-head attention mask module is employed to learn the soft attention mask features from the skeletons, suppressing non-salient local features while accentuating local salient features, thereby bringing similar local features closer in the feature space. Additionally, ample contrastive pairs are generated by expanding contrastive pairs based on salient and non-salient features with global features, which guide the network to learn the semantic representations of the entire skeleton. Therefore, with the attention mask mechanism, SkeAttnCLR learns local features under different data augmentation views. The experiment results demonstrate that the inclusion of local feature similarity significantly enhances skeleton-based action representation. Our proposed SkeAttnCLR outperforms state-of-the-art methods on NTURGB+D, NTU120-RGB+D, and PKU-MMD datasets. The code and settings are available at this repository: https://github.com/GitHubOfHyl97/SkeAttnCLR.