Generating Human-Like Behaviors Using Joint, Speech-Driven Models for Conversational Agents

Generating Human-Like Behaviors Using Joint, Speech-Driven Models for Conversational Agents
复制标题

使用会话代理的联合语音驱动模型生成类人行为

DOI:
10.1109/tasl.2012.2201476
复制
发表时间:
2012
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
C. Busso
C. Busso
中科院分区:
--
文献类型:
--
作者:
Soroosh Mariooryad;C. Busso

文献摘要

被引文献

相似文献

在人类交流过程中,每一个口头信息都在不同的语言和非语言线索中内在地调制,这些线索通过言语和面部表情的各个方面外在化。这些沟通渠道是密切相关的,这表明产生类似人类的行为需要仔细研究它们之间的关系。在对会话代理的自然行为进行建模时忽略不同通信通道的相互影响可能会导致不切实际的行为,这些行为会影响动画的预期视觉感知。这种关系既存在于视听信息之间,也存在于不同的视觉方面。本文探讨的想法,使用联合模型,以保持不仅在语音和面部表情之间的耦合,但也在面部表情。作为一个案例研究,本文重点是建立一个语音驱动的人脸动画框架,以产生自然的头部和眉毛的运动。我们提出了三个动态贝叶斯网络(DBNs),这使得不同的假设之间的耦合语音,眉毛和头部运动。使用视听IEMOCAP数据库,根据MPEG-4面部动画标准制作合成动画。基于感知评估的实验结果表明,所提出的联合模型(语音/眉毛/头部)优于单独训练的视听模型(语音/头部和语音/眉毛)。
During human communication, every spoken message is intrinsically modulated within different verbal and nonverbal cues that are externalized through various aspects of speech and facial gestures. These communication channels are strongly interrelated, which suggests that generating human-like behavior requires a careful study of their relationship. Neglecting the mutual influence of different communicative channels in the modeling of natural behavior for a conversational agent may result in unrealistic behaviors that can affect the intended visual perception of the animation. This relationship exists both between audiovisual information and within different visual aspects. This paper explores the idea of using joint models to preserve the coupling not only between speech and facial expression, but also within facial gestures. As a case study, the paper focuses on building a speech-driven facial animation framework to generate natural head and eyebrow motions. We propose three dynamic Bayesian networks (DBNs), which make different assumptions about the coupling between speech, eyebrow and head motion. Synthesized animations are produced based on the MPEG-4 facial animation standard, using the audiovisual IEMOCAP database. The experimental results based on perceptual evaluations reveal that the proposed joint models (speech/eyebrow/head) outperform audiovisual models that are separately trained (speech/head and speech/eyebrow).