Skeleton-Based Human Action Recognition With Global Context-Aware Attention LSTM Networks

Skeleton-Based Human Action Recognition With Global Context-Aware Attention LSTM Networks
复制标题

DOI:
10.1109/tip.2017.2785279
复制
发表时间:
2018-04-01
影响因子:
10.6
通讯作者:
Kot, Alex C.
Kot, Alex C.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Liu, Jun;Wang, Gang;Kot, Alex C.

文献摘要

被引文献

相似文献

三维人体骨架序列中的人体动作识别已经引起了人们的广泛关注。最近,长短期记忆(LSTM)网络在这一任务中表现出了良好的性能,因为它们在建模序列数据中的依赖性和动态方面具有优势。由于并不是所有的骨骼关节都能为动作识别提供信息,并且不相关的关节往往会带来噪声,从而降低性能,因此我们需要更多地关注信息关节。然而,原始的LSTM网络不具有显式注意力能力。在本文中,我们提出了一类新的LSTM网络,全局上下文感知注意力LSTM,用于基于动作的动作识别,它能够通过使用全局上下文记忆单元来选择性地关注每帧中的信息关节。为了进一步提高网络的注意力,我们引入了一种循环注意机制,通过这种机制,我们的网络的注意力性能可以逐步提高。此外,本文还提出了一个利用粗粒度注意力和细粒度注意力的双流框架。所提出的方法在五个具有挑战性的数据集上实现了最先进的性能,用于基于动作的动作识别。
Human action recognition in 3D skeleton sequences has attracted a lot of research attention. Recently, long short-term memory (LSTM) networks have shown promising performance in this task due to their strengths in modeling the dependencies and dynamics in sequential data. As not all skeletal joints are informative for action recognition, and the irrelevant joints often bring noise which can degrade the performance, we need to pay more attention to the informative ones. However, the original LSTM network does not have explicit attention ability. In this paper, we propose a new class of LSTM network, global context-aware attention LSTM, for skeleton-based action recognition, which is capable of selectively focusing on the informative joints in each frame by using a global context memory cell. To further improve the attention capability, we also introduce a recurrent attention mechanism, with which the attention performance of our network can be enhanced progressively. Besides, a two-stream framework, which leverages coarse-grained attention and fine-grained attention, is also introduced. The proposed method achieves state-of-the-art performance on five challenging datasets for skeleton-based action recognition.