VISTURE: A System for Video-Based Gesture and Speech Generation by Robots

VISTURE: A System for Video-Based Gesture and Speech Generation by Robots
复制标题

DOI:
10.1145/3527188.3561931
复制
发表时间:
2022-12
期刊:
Proceedings of the 10th International Conference on Human-Agent Interaction
影响因子:
--
通讯作者:
Kaon Shimoyama;Kohei Okuoka;Mitsuhiko Kimoto;M. Imai
Kaon Shimoyama;Kohei Okuoka;Mitsuhiko Kimoto;M. Imai
中科院分区:
其他
文献类型:
--
作者:
Kaon Shimoyama;Kohei Okuoka;Mitsuhiko Kimoto;M. Imai

文献摘要

相似文献

本文提出了一种以视频为输入的机器人手势和语音生成系统VISTURE。VISTURE假设了这样一种情况,即机器人将它用相机看到的东西传达给一个缺席的人。本文的价值在于,我们对日本人用来描述视频场景的表达方式进行了个案研究,并将研究结果用于构建VISTURE。特别是,我们在整个案例研究中发现了描述视频场景的表情的分类:前景信息是场景的相关事件,背景信息不是给出整个场景的描述的重点。前景和背景是组合指代的。VISTURE使用分类来生成类似人类的表情。此外,我们设计了确定前景和背景的方法,该方法可以生成多个表情组合。我们调查了人们对机器人执行VISTURE生成的手势和语音的印象,以评估这些手势和语音的质量。结果显示,机器人在做手势时更讨人喜欢,也更能干。
This paper proposes VISTURE, a system for generating a robot’s gesture and speech by using video as input. VISTURE assumes a situation in which a robot conveys what it saw with a camera to a person who was absent. The value of this paper is that we have performed a case study to investigate the expressions that Japanese people use to describe video scenes, and used the results to build VISTURE. In particular, we found classification of expressions depicting the video scenes throughout the case study: Foreground information that is the relevant event of the scene and Background one that is not the main point of the description giving the entire scene. Foreground and Background are referred in combination. VISTURE employs the classification to generate human-like expressions. Moreover, we designed the method to determine Foreground and Background, and it can generate multiple combinations of expressions. We investigated the people’s impression of a robot performing the gestures and speech generated by VISTURE to evaluate the quality of those gestures and speech. The results showed that the robot was perceived as more likable and capable when it performed gestures.