AL-SAR: Active Learning for Skeleton-Based Action Recognition

AL-SAR: Active Learning for Skeleton-Based Action Recognition
复制标题

AL-SAR:基于骨架的动作识别的主动学习

DOI:
10.1109/tnnls.2023.3297853
复制
发表时间:
2023
影响因子:
10.4
通讯作者:
Shlizerman, Eli
Shlizerman, Eli
中科院分区:
计算机科学1区
文献类型:
--
作者:
Li, Jingyuan;Le, Trung;Shlizerman, Eli

文献摘要

相似文献

从时间多变量特征序列中进行动作识别,例如识别人类动作,通常通过监督训练来实现,因为它需要许多基本事实注释才能达到高识别精度。已经引入了用于将序列组织成簇的非监督方法,然而,这种方法继续需要注释来将簇与动作相关联。注释方面的挑战需要一种有效的分类方法,将所需的标签数量降至最低。主动学习(AL)方法已经被提出来解决这些挑战,并能够在图像分类上建立稳健的结果。这种方法不能直接应用于序列,因为对于序列,变化既在空间域又在时间域。在本文中,我们介绍了一种新的序列人工智能方法,称为AL-SAR,它结合了无监督训练和稀疏监督标注。特别是,AL-SAR采用多头机制对编解码器框架学习的潜在空间进行稳健的不确定性评估。它的目标是迭代地选择一组稀疏的样本,其中的注释对潜在空间的解缠贡献最大。我们在具有多个序列和动作的公共基准数据集上对我们的系统进行了评估,例如NW-UCLA、NTU RGB+D60和UWA3D。结果表明,与编解码器网络相耦合的AL-SAR的性能优于相同网络结构的其他AL方法。
Action recognition from temporal multivariate sequences of features, such as identifying human actions, is typically approached by supervised training as it requires many ground truth annotations to reach high recognition accuracy. Unsupervised methods for the organization of sequences into clusters have been introduced, however, such methods continue to require annotations to associate clusters with actions. The challenges in annotation necessitate an effective classification methodology that minimizes the required number of labels. Active learning (AL) approaches have been proposed to address these challenges and were able to establish robust results on image classification. Such approaches are not directly applicable to sequences, since for sequences, the variations are in both spatial and temporal domains. In this brief, we introduce a novel method for AL for sequences, called “AL-SAR,” which combines unsupervised training with sparsely supervised annotation. In particular, AL-SAR employs a multi-head mechanism for robust uncertainty evaluation of the latent space learned by an encoder-decoder framework. It aims to iteratively select a sparse set of samples, which annotation contributes the most to the disentanglement of the latent space. We evaluate our system on common benchmark datasets with multiple sequences and actions, such as NW-UCLA, NTU RGB+D 60, and UWA3D. Our results indicate that AL-SAR coupled with encoder-decoder network outperforms other AL methods coupled with the same network structure.