Learning Backchanneling Behaviors for a Social Robot via Data Augmentation from Human-Human Conversations

Learning Backchanneling Behaviors for a Social Robot via Data Augmentation from Human-Human Conversations
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Michael Murray;Nick Walker;Amal Nanavati;Patrícia Alves-Oliveira;Nikita Filippov;Allison Sauppé;Bilge Mutlu;M. Cakmak
Michael Murray;Nick Walker;Amal Nanavati;Patrícia Alves-Oliveira;Nikita Filippov;Allison Sauppé;Bilge Mutlu;M. Cakmak
中科院分区:
其他
文献类型:
--
作者:
Michael Murray;Nick Walker;Amal Nanavati;Patrícia Alves-Oliveira;Nikita Filippov;Allison Sauppé;Bilge Mutlu;M. Cakmak

文献摘要

相似文献

机器人的反向行为,例如点头,可以让机器人感觉到机器人正在积极倾听,从而使与机器人交谈变得更加自然和吸引人。为了使反向沟通有效,重要的是,这种暗示的时机是适当的,考虑到人类的会话行为。最近的进展表明,这些行为可以从人与人之间的对话数据集中学习。然而,最近的数据驱动方法往往过度拟合训练数据中看到的人类说话者,并且无法很好地推广到以前看不见的说话者。在本文中,我们探讨了使用数据增强有效点头行为的机器人。我们表明,通过增强输入语音和视觉特征,我们可以生成数据驱动的模型,这些模型对看不见的特征更鲁棒,而无需收集额外的数据。我们分析了数据驱动的反向通道在现实的人机对话环境中的有效性,并进行了用户研究,结果表明,与基于规则和随机基线相比,用户认为数据驱动的模型更善于倾听。
: Backchanneling behaviors on a robot, such as nodding, can make talking to a robot feel more natural and engaging by giving a sense that the robot is actively listening. For backchanneling to be effective, it is important that the timing of such cues is appropriate given the humans’ conversational behaviors. Recent progress has shown that these behaviors can be learned from datasets of human-human conversations. However, recent data-driven methods tend to overfit to the human speakers that are seen in training data and fail to generalize well to previously unseen speakers. In this paper, we explore the use of data augmentation for effective nodding behavior in a robot. We show that, by augmenting the input speech and visual features, we can produce data-driven models that are more robust to unseen features without collecting additional data. We analyze the efficacy of data-driven backchanneling in a realistic human-robot conversational setting with a user study, showing that users perceived the data-driven model to be better at listening as compared to rule-based and random baselines.