Generation of action description from classification of motion and object

Generation of action description from classification of motion and object
复制标题

DOI:
10.1016/j.robot.2017.02.003
复制
发表时间:
2017-05
期刊:
Robotics Auton. Syst.
影响因子:
--
通讯作者:
W. Takano;Y. Yamada;Yoshihiko Nakamura
W. Takano;Y. Yamada;Yoshihiko Nakamura
中科院分区:
其他
文献类型:
--
作者:
W. Takano;Y. Yamada;Yoshihiko Nakamura

文献摘要

相似文献

本文提出了一种新的方法来学习运动、物体和语言之间的关系,并生成描述人类行为的句子。我们的方法对人体运动和作用于这些运动的物体进行分类,然后将运动类别和物体类别与它们的描述性句子相结合。集成包括两个步骤。第一步随机学习句子中动作、物体和单词之间的关系。第二步随机学习句子中单词的顺序作为句子的结构。第一步导出的模型称为“动作语言”模型,第二步导出的模型称为“自然语言”模型。这种将动作语言模型与自然语言模型相结合的框架可以应用于从人类动作中生成描述性句子,其中每个动作被识别为包含动作类别和对象类别的一对,通过包含的动作和对象类别生成与动作相关的单词,并将单词排列成描述性句子。从理论上讲,我们的方法通过使用动作语言模型来搜索可能从动作和对象类别中生成的多个单词;然后使用自然语言模型搜索这些词的序列,这些词很可能是从获得的词中生成的。我们通过将该方法应用于RGB-D传感器捕获的人类动作数据来测试该方法的句子生成,并证明了其有效性。
This paper presents a novel approach to learning of relations among motions, objects, and language, and to generating sentences that describe human actions. Our approach categorizes human motions and the objects acted on those motions, and subsequently integrates the motion categories and object categories with their descriptive sentences. The integration consists of two steps. The first step stochastically learns the relations among the motions, objects, and words in the sentences. The second step stochastically learns the order of words in the sentences as the sentence structures. The model derived in the first step is referred to as “action language” model and that derived in the second step as “natural language” model. This framework for integrating an action language model with a natural language model can be applied to generating descriptive sentences from human actions, where each action is recognized as a pair containing a motion category and an object category, the words relevant to the action are generated via the contained motion and object categories, and the words to be arranged result in a descriptive sentence. More theoretically, our approach searches for multiple words likely to be generated from the motion and object categories by using the action language model; and subsequently searches for a sequence of these words that is likely to be generated from the obtained words, using the natural language model. We tested our proposed approach for sentence generation by applying it to human action data captured by an RGB-D sensor, and demonstrated its validity.