Enhancing human action recognition through spatio-temporal feature learning and semantic rules

Enhancing human action recognition through spatio-temporal feature learning and semantic rules
复制标题

通过时空特征学习和语义规则增强人类动作识别

DOI:
--
复制
发表时间:
2013
期刊:
IEEE-RAS International Conference on Humanoid Robots
影响因子:
--
通讯作者:
G. Cheng
G. Cheng
中科院分区:
--
文献类型:
--
作者:
Karinne Ramirez;Eun;Jiseob Kim;Byoung;M. Beetz;G. Cheng

文献摘要

被引文献

相似文献

在本文中,我们提出了一个两阶段的框架来处理从视频中自动提取人类活动的问题。首先,对于动作识别,我们采用了一种基于独立子空间分析(ISA)的无监督学习算法。该学习算法直接从视频数据中提取时空特征,与其他非监督方法相比,具有更高的计算效率和健壮性。然而,当将这一阶段最先进的动作识别技术应用于人类日常活动的观察时,其准确率只能达到25%左右。因此,我们建议用第二阶段来增强这一过程,该阶段定义了一种新的方法来自动生成能够对人类活动进行推理的语义规则。得到的语义规则降低了感知系统的复杂性,增强了对人类行为的识别能力,并允许领域变化的可能性,从而极大地提高了机器人行为的综合能力。所提出的方法在两个复杂和具有挑战性的场景下进行了评估:制作煎饼和制作三明治。这些场景的困难在于,它们包含比众所周知的数据集(Hollywood 2、KTH等)更精细、更复杂的活动。实验结果表明,两阶段方法的优点是,动作识别的准确率比单阶段方法有了显著提高(与人类专家相比提高了87%以上)。这表明使用推理引擎从观测中自动提取人类活动的框架得到了改进,从而为将广泛的人类技能转移到类人机器人提供了丰富的机制。
In this paper, we present a two-stage framework that deal with the problem of automatically extract human activities from videos. First, for action recognition we employ an unsupervised state-of-the-art learning algorithm based on Independent Subspace Analysis (ISA). This learning algorithm extracts spatio-temporal features directly from video data and it is computationally more efficient and robust than other unsupervised methods. Nevertheless, when applying this one-stage state-of-the-art action recognition technique on the observations of human everyday activities, it can only reach an accuracy rate of approximately 25%. Hence, we propose to enhance this process with a second stage, which define a new method to automatically generate semantic rules that can reason about human activities. The obtained semantic rules enhance the human activity recognition by reducing the complexity of the perception system and they allow the possibility of domain change, which can great improve the synthesis of robot behaviors. The proposed method was evaluated under two complex and challenging scenarios: making a pancake and making a sandwich. The difficulty of these scenarios is that they contain finer and more complex activities than the well known data sets (Hollywood2, KTH, etc). The results show benefits of two stages method, the accuracy of action recognition was significantly improved compared to a single-stage method (above 87% compared to human expert). This indicates the improvement of the framework using the reasoning engine for the automatic extraction of human activities from observations, thus, providing a rich mechanism for transferring a wide range of human skills to humanoid robots.