Early Recognition of 3D Human Actions

Early Recognition of 3D Human Actions
复制标题

DOI:
10.1145/3131344
复制
发表时间:
2018-03
期刊:
ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)
影响因子:
--
通讯作者:
Sheng Li;Kang Li;Y. Fu
Sheng Li;Kang Li;Y. Fu
中科院分区:
其他
文献类型:
--
作者:
Sheng Li;Kang Li;Y. Fu

文献摘要

被引文献

相似文献

动作识别是人体运动分析的一个重要研究课题。近年来,基于3D观察的动作识别在多媒体和计算机视觉社区中受到越来越多的关注,这是由于最近出现了具有成本效益的传感器,例如深度相机Kinect。这项工作更进一步,专注于早期识别正在进行的3D人类行为,这对于各种时间关键型应用程序都是有益的,例如,基于手势的人机交互、体感游戏等等。我们的目标是推断类标签信息的3D人类行动的部分观察时间不完整的行动执行。通过将3D动作数据视为多变量时间序列(m.t.s.)与共享的公共时钟(帧)同步,我们提出了一种称为动态标记点过程(DMP)的随机过程,将3D动作建模为时间动态模式,其中捕获时间和强度信息。为了实现更早和更好的识别精度,我们还探索了特征维度之间的时间依赖模式。基于变阶马尔可夫模型,构造了一棵概率后缀树来表示特征间的序列模式。我们的方法和几个基线在五个3D人体动作数据集上进行了评估。大量的结果表明,我们的方法实现了上级性能的早期识别的三维人体动作。
Action recognition is an important research problem of human motion analysis (HMA). In recent years, 3D observation-based action recognition has been receiving increasing interest in the multimedia and computer vision communities, due to the recent advent of cost-effective sensors, such as depth camera Kinect. This work takes this one step further, focusing on early recognition of ongoing 3D human actions, which is beneficial for a large variety of time-critical applications, e.g., gesture-based human machine interaction, somatosensory games, and so forth. Our goal is to infer the class label information of 3D human actions with partial observation of temporally incomplete action executions. By considering 3D action data as multivariate time series (m.t.s.) synchronized to a shared common clock (frames), we propose a stochastic process called dynamic marked point process (DMP) to model the 3D action as temporal dynamic patterns, where both timing and strength information are captured. To achieve even more early and better accuracy of recognition, we also explore the temporal dependency patterns between feature dimensions. A probabilistic suffix tree is constructed to represent sequential patterns among features in terms of the variable-order Markov model (VMM). Our approach and several baselines are evaluated on five 3D human action datasets. Extensive results show that our approach achieves superior performance for early recognition of 3D human actions.