NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding

NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding
复制标题

DOI:
10.1109/tpami.2019.2916873
复制
发表时间:
2020-10-01
影响因子:
23.6
通讯作者:
Kot, Alex C.
Kot, Alex C.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Liu, Jun;Shahroudy, Amir;Kot, Alex C.

文献摘要

被引文献

相似文献

基于深度的人体活动分析的研究取得了出色的性能,并证明了3D表示的动作识别的有效性。现有的基于深度和基于RGB+ D的动作识别基准具有许多局限性,包括缺乏大规模训练样本,不同类别类别的实际数量,相机视图的多样性,各种环境条件以及各种人类主体。在这项工作中,我们介绍了一个用于RGB+D人体动作识别的大规模数据集,该数据集从106个不同的主题中收集,包含超过11.4万个视频样本和800万帧。该数据集包含120个不同的动作类,包括日常、相互和健康相关的活动。我们评估了一系列现有的3D活动分析方法在该数据集上的性能,并展示了将深度学习方法应用于基于3D的人类动作识别的优势。此外,我们研究了一个新的一次性3D活动识别问题,我们的数据集,并提出了一个简单而有效的部分语义相关性感知(APSR)框架,这产生了有前途的结果识别的新的动作类。我们相信,这个大规模数据集的引入将使社区能够应用、适应和开发各种数据饥渴的学习技术,用于基于深度和RGB+ D的人类活动理解。
Research on depth-based human activity analysis achieved outstanding performance and demonstrated the effectiveness of 3D representation for action recognition. The existing depth-based and RGB+D-based action recognition benchmarks have a number of limitations, including the lack of large-scale training samples, realistic number of distinct class categories, diversity in camera views, varied environmental conditions, and variety of human subjects. In this work, we introduce a large-scale dataset for RGB+D human action recognition, which is collected from 106 distinct subjects and contains more than 114 thousand video samples and 8 million frames. This dataset contains 120 different action classes including daily, mutual, and health-related activities. We evaluate the performance of a series of existing 3D activity analysis methods on this dataset, and show the advantage of applying deep learning methods for 3D-based human action recognition. Furthermore, we investigate a novel one-shot 3D activity recognition problem on our dataset, and a simple yet effective Action-Part Semantic Relevance-aware (APSR) framework is proposed for this task, which yields promising results for recognition of the novel action classes. We believe the introduction of this large-scale dataset will enable the community to apply, adapt, and develop various data-hungry learning techniques for depth-based and RGB+D-based human activity understanding.