Data-driven synthesis of expressive visual speech using an MPEG-4 talking head

Data-driven synthesis of expressive visual speech using an MPEG-4 talking head
复制标题

使用 MPEG-4 说话头进行数据驱动的富有表现力的视觉语音合成

DOI:
--
复制
发表时间:
2005
期刊:
Interspeech
影响因子:
--
通讯作者:
M. Nordenberg
M. Nordenberg
中科院分区:
--
文献类型:
--
作者:
J. Beskow;M. Nordenberg

文献摘要

被引文献

相似文献

本文介绍了初步的实验与合成的视觉语音清晰度不同的情绪,使用新开发的MPEG-4兼容的说话头。在说话的头部中结合语音和情感的基本问题是处理口面区域中情感表达和清晰度之间的相互作用。这里采取的策略不是试图将言语和情感建模为两个独立的属性,而是从一开始就将情感表达纳入发音中。我们使用数据驱动的方法,训练系统重新创建演员在描绘不同情绪时产生的表达性清晰度。每种情绪分别使用主成分分析和参数协同表达模型建模。到目前为止的结果是令人鼓舞的,但需要更多的工作来提高合成语音的自然度和准确性。
This paper describes initial experiments with synthesis of visual speech articulation for different emotions, using a newly developed MPEG-4 compatible talking head. The basic problem with combining speech and emotion in a talking head is to handle the interaction between emotional expression and articulation in the orofacial region. Rather than trying to model speech and emotion as two separate properties, the strategy taken here is to incorporate emotional expression in the articulation from the beginning. We use a data-driven approach, training the system to recreate the expressive articulation produced by an actor while portraying different emotions. Each emotion is modelled separately using principal component analysis and a parametric coarticulation model. The results so far are encouraging but more work is needed to improve naturalness and accuracy of the synthesized speech.