A Review of Evaluation Practices of Gesture Generation in Embodied Conversational Agents

A Review of Evaluation Practices of Gesture Generation in Embodied Conversational Agents
复制标题

具身对话代理中手势生成的评估实践综述

DOI:
10.1109/thms.2022.3149173
复制
发表时间:
2021
影响因子:
3.6
通讯作者:
Tony Belpaeme
Tony Belpaeme
中科院分区:
计算机科学3区
文献类型:
--
作者:
Pieter Wolfert;Nicole L. Robinson;Tony Belpaeme

文献摘要

参考文献

被引文献

相似文献

具身会话代理(ECA)通常被设计为产生非语言行为,以补充或增强他们的语言交流。非语言行为的一种形式是协同语音手势,它涉及代理用手臂和手做出的与语言交流相匹配的动作。 ECA 的协同语音手势可以使用不同的生成方法创建,分为基于规则的过程和数据驱动的过程,后者由于应用机器学习社区的兴趣日益浓厚而受到关注。然而,关于手势生成方法的报告使用了多种评估措施,这阻碍了比较。为了解决这个问题,我们对标志性、隐喻性、指示性和节拍手势的协同语音手势生成方法进行了系统回顾,包括报告的评估方法。我们回顾了 22 项研究,这些研究的 ECA 具有类似人类的上半身,在社交人机交互中使用共同语音手势。这包括使用人类参与者来评估表现的研究。我们发现大多数研究都采用受试者内设计并依赖于某种形式的主观评估,但没有系统的方法。我们认为,该领域需要更严格和统一的工具来进行协同语音手势评估,并制定实证评估建议,包括标准化短语和示例场景,以帮助系统地测试跨研究的生成模型。此外,我们还提出了一个清单,可用于报告生成模型评估的相关信息,以及评估协同语音手势的使用。
Embodied conversational agents (ECAs) are often designed to produce nonverbal behavior to complement or enhance their verbal communication. One such form of the nonverbal behavior is co-speech gesturing, which involves movements that the agent makes with its arms and hands that are paired with verbal communication. Co-speech gestures for ECAs can be created using different generation methods, divided into rule-based and data-driven processes, with the latter, gaining traction because of the increasing interest from the applied machine learning community. However, reports on gesture generation methods use a variety of evaluation measures, which hinders comparison. To address this, we present a systematic review on co-speech gesture generation methods for iconic, metaphoric, deictic, and beat gestures, including reported evaluation methods. We review 22 studies that have an ECA with a human-like upper body that uses co-speech gesturing in social human-agent interaction. This includes studies that use human participants to evaluate performance. We found most studies use a within-subject design and rely on a form of subjective evaluation, but without a systematic approach. We argue that the field requires more rigorous and uniform tools for co-speech gesture evaluation, and formulate recommendations for empirical evaluation, including standardized phrases and example scenarios to help systematically test generative models across studies. Furthermore, we also propose a checklist that can be used to report relevant information for the evaluation of generative models, as well as to evaluate co-speech gesture use.
DOI: 10.1007/s12369-013-0196-9
发表时间: 2013-08-01
影响因子: 4.7
作者:
Salem, Maha;Eyssel, Friederike;Joublin, Frank
通讯作者: Joublin, Frank
机器学习和数据挖掘百科全书
DOI: 10.1007/978-1-4899-7502-7_900-1
发表时间: 2016
期刊: --
影响因子: --
作者:
Flach P
通讯作者: Flach P