Agree or Disagreeƒ Generating Body Gestures from Affective Contextual Cues during Dyadic Interactions

Agree or Disagreeƒ Generating Body Gestures from Affective Contextual Cues during Dyadic Interactions
复制标题

同意或不同意在二元交互过程中根据情感上下文线索生成身体手势

DOI:
--
复制
发表时间:
2022
期刊:
IEEE International Symposium on Robot and Human Interactive Communication
影响因子:
--
通讯作者:
O. Çeliktutan
O. Çeliktutan
中科院分区:
--
文献类型:
--
作者:
Nguyen Tan Viet Tuyen;O. Çeliktutan

文献摘要

被引文献

相似文献

人类自然会产生非语言信号,如面部表情、身体动作、手势和语调,以及文字,来传达他们的信息、观点和感受。考虑到机器人正逐渐从研究实验室走向人类环境,人们越来越希望它们发展出类似的社交智能。因此,几十年来,为社交机器人配备非语言沟通技能一直是一个活跃的研究领域,近年来,数据驱动的端到端学习方法已成为主导,具有可扩展性和通用性。然而,这些方法大多只考虑单个角色,只模拟个人动态。在本文中,我们提出了一种基于条件生成对抗网络的方法,旨在为机器人在情感二元交互中生成行为。我们的方法将目标人的音频和他们互动伙伴的非语言信号作为输入,通过一种新的上下文编码器建模,生成适当的肢体动作。我们在多模态JESTKOD数据集上评估了我们的方法,该数据集包括一致和不一致场景下的二元交互。实验结果表明,上下文编码器能够更好地预测协议情境下的同语手势。
Humans naturally produce nonverbal signals such as facial expressions, body movements, hand gestures, and tone of voice, along with words, to communicate their messages, opinions, and feelings. Considering robots are progressively moving out from research laboratories into human environments, it is increasingly desirable that they develop a similar social intelligence. Therefore, equipping social robots with nonverbal communication skills has been an active research area for decades, where data-driven, end-to-end learning approaches have become predominant in recent years, offering scalability and generalisability. However, most of these approaches consider a single character, modelling intrapersonal dynamics only. In this paper, we propose a method based on conditional Generative Adversarial Networks, intending to generate behaviours for a robot in affective dyadic interactions. Our method takes as an input the audio of a target person together with the nonverbal signals of their interacting partner, modelled by a novel Context Encoder, to generate appropriate body gestures. We evaluate our method on the multimodal JESTKOD dataset that comprises dyadic interactions under agreement and disagreement scenarios. The experimental results show that Context Encoder can better contribute to the prediction of co-speech gestures in agreement situations.