Learning Expressive Human-Like Head Motion Sequences from Speech

Learning Expressive Human-Like Head Motion Sequences from Speech
复制标题

从语音中学习富有表现力的类人头部运动序列

DOI:
--
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
Shrikanth S. Narayanan
Shrikanth S. Narayanan
中科院分区:
--
文献类型:
--
作者:
C. Busso;Z. Deng;U. Neumann;Shrikanth S. Narayanan

文献摘要

被引文献

相似文献

随着人机界面、动画故事片和视频游戏等新趋势的发展,人们需要更好的化身和虚拟代理来更准确地模拟人类的交流和互动方式。手势和言语共同用来表达预期的信息。讲话的语气和能量、面部表情、僵硬的头部动作和手部动作以一种不平凡的方式结合在一起,因为它们在自然的人类互动中展开。考虑到使用大型运动捕捉数据集的成本很高,并且只能应用于计划的场景中,因此需要新的自动方法来合成逼真的动画,以捕捉并类似于这些交流通道之间的复杂关系。一种有用和实用的方法是使用声学特征来生成手势,利用手势和语音之间的联系。由于嘴唇的形状由潜在的发音决定,声学特征被用来生成与口语句子匹配的视觉视位[4,5,12,17]。同样,声学特征也被用来合成面部表情[11,30],利用了用于发音的相同肌肉也会影响脸型的事实[44,46]。在面部动画中,与其他方面相比,一个重要的手势受到的关注较少,那就是刚性头部运动。头部运动不仅对于确认积极倾听或取代言语信息(例如“点头”)很重要,而且对人类的许多方面也是重要的
With the development of new trends in human-machine interfaces, animated feature films and video games, better avatars and virtual agents are required that more accurately mimic how humans communicate and interact. Gestures and speech are jointly used to express intended messages. The tone and energy of the speech, facial expression, rigid head motion and hand motion combine in a non-trivial manner as they unfold in natural human interaction. Given that the use of large motion capture datasets is expensive and can only be applied in planned scenarios, new automatic approaches are required to synthesize realistic animation that capture and resemble the complex relationship between these communicative channels. One useful and practical approach is the use of acoustic features to generate gestures, exploiting the link between gestures and speech. Since the shape of the lips is determined by the underlying articulation, acoustic features have been used to generate visual visemes that match the spoken sentences [4, 5, 12, 17]. Likewise, acoustic features have been used to synthesize facial expressions [11, 30], exploiting the fact that the same muscles used for articulation also affect the shape of the face [44, 46]. One important gesture that has received less attention than other aspects in facial animations is rigid head motion. Head motion is important not only to acknowledge active listening or replace verbal information (e.g. “nod”), but also for many aspect of human