Linking human motions and objects to language for synthesizing action sentences

Linking human motions and objects to language for synthesizing action sentences
复制标题

DOI:
10.1007/s10514-018-9762-1
复制
发表时间:
2018-05
期刊:
影响因子:
3.5
通讯作者:
W. Takano;Y. Yamada;Yoshihiko Nakamura
W. Takano;Y. Yamada;Yoshihiko Nakamura
中科院分区:
计算机科学3区
文献类型:
--
作者:
W. Takano;Y. Yamada;Yoshihiko Nakamura

文献摘要

相似文献

本文提出了一种新的框架,从人体全身运动和对象的操作产生的动作描述。这种生成基于三个模块:第一模块对人类运动和对象进行分类;第二模块将运动和对象类别与单词相关联;第三模块提取句子结构作为单词序列。首先在第一模块中对人体运动和待操作对象进行分类,然后在第二模块中生成与运动和对象类别高度相关的单词,最后在第三模块中将单词转换为单词序列形式的句子。运动和对象沿着运动、对象和词之间的关系由第一和第二模块随机参数化。第三模块从动态系统中的词序列数据集参数化句子结构。运动、物体和词语的随机表示与句子的动态表示的链接允许合成描述人类动作的句子。我们测试了我们所提出的方法合成的动作描述的RGB-D传感器捕获的人类动作数据集,并证明了其有效性。
This paper proposes a novel framework for generating action descriptions from human whole body motions and objects to be manipulated. This generation is based on three modules: the first module categorizes human motions and objects; the second module associates the motion and object categories with words; and the third module extracts a sentence structure as word sequences. Human motions and objects to be manipulated are classified into categories in the first module, then words highly relevant to the motion and object categories are generated from the second module, and finally the words are converted into sentences in the form of word sequences by the third module. The motions and objects along with the relations among the motions, objects, and words are parametrized stochastically by the first and second modules. The sentence structures are parametrized from a dataset of word sequences in a dynamical system by the third module. The link of the stochastic representation of the motions, objects, and words with the dynamical representation of the sentences allows for synthesizing sentences descriptive to human actions. We tested our proposed method on synthesizing action descriptions for a human action dataset captured by an RGB-D sensor, and demonstrated its validity.