Data-driven synthesis of expressive visual speech using an MPEG-4 talking head
Data-driven synthesis of expressive visual speech using an MPEG-4 talking head
复制标题
使用 MPEG-4 说话头进行数据驱动的富有表现力的视觉语音合成
DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
M. Nordenberg
中科院分区:
文献类型:
--
作者:
J. Beskow;M. Nordenberg
This paper describes initial experiments with synthesis of visual speech articulation for different emotions, using a newly developed MPEG-4 compatible talking head. The basic problem with combining speech and emotion in a talking head is to handle the interaction between emotional expression and articulation in the orofacial region. Rather than trying to model speech and emotion as two separate properties, the strategy taken here is to incorporate emotional expression in the articulation from the beginning. We use a data-driven approach, training the system to recreate the expressive articulation produced by an actor while portraying different emotions. Each emotion is modelled separately using principal component analysis and a parametric coarticulation model. The results so far are encouraging but more work is needed to improve naturalness and accuracy of the synthesized speech.