Corpus-based generation of head and eyebrow motion for an embodied conversational agent

Corpus-based generation of head and eyebrow motion for an embodied conversational agent
复制标题

基于语料库的具体对话代理的头部和眉毛运动生成

DOI:
--
复制
发表时间:
2007
影响因子:
2.7
通讯作者:
J. Oberlander
J. Oberlander
中科院分区:
计算机科学4区
文献类型:
--
作者:
Mary Ellen Foster;J. Oberlander

文献摘要

被引文献

相似文献

众所周知,人类在说话时会使用各种各样的非语言行为。因此,为人工智能体生成自然的具体化语音是一种应用,在这种应用中,直接利用记录的人类动作的技术可能会有所帮助。我们提出了一个系统,该系统使用基于语料库的选择策略来指定动画说话头的头部和眉毛的运动。我们首先描述了如何记录和注释特定领域的面部显示语料库,并概述了在数据中发现的规律。然后,我们根据语料库数据提出了两种不同的选择说话头动作的方法:一种是在所有情况下选择多数选项,另一种是在所有选项中进行加权选择。我们通过两种方式对这些方法进行比较:通过对语料库的交叉验证,并要求人类法官对输出进行评分。两个评估研究的结果不同:交叉验证研究倾向于多数策略,而人类法官更喜欢使用加权选择生成的时间表。在第二项研究中,评委也表现出对原始语料库数据的偏好,而不是任何一种生成策略的输出。
Humans are known to use a wide range of non-verbal behaviour while speaking. Generating naturalistic embodied speech for an artificial agent is therefore an application where techniques that draw directly on recorded human motions can be helpful. We present a system that uses corpus-based selection strategies to specify the head and eyebrow motion of an animated talking head. We first describe how a domain-specific corpus of facial displays was recorded and annotated, and outline the regularities that were found in the data. We then present two different methods of selecting motions for the talking head based on the corpus data: one that chooses the majority option in all cases, and one that makes a weighted choice among all of the options. We compare these methods to each other in two ways: through cross-validation against the corpus, and by asking human judges to rate the output. The results of the two evaluation studies differ: the cross-validation study favoured the majority strategy, while the human judges preferred schedules generated using weighted choice. The judges in the second study also showed a preference for the original corpus data over the output of either of the generation strategies.