Realistic Mouth-Synching for Speech-Driven Talking Face Using Articulatory Modelling

Realistic Mouth-Synching for Speech-Driven Talking Face Using Articulatory Modelling
复制标题

DOI:
10.1109/tmm.2006.888009
复制
发表时间:
2007-04
影响因子:
7.3
通讯作者:
Lei Xie;Zhi-Qiang Liu
Lei Xie;Zhi-Qiang Liu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Lei Xie;Zhi-Qiang Liu

文献摘要

被引文献

相似文献

本文提出了一种将声学语音转换为逼真的嘴巴动画的发音建模方法。我们使用基于动态贝叶斯网络 (DBN) 的视听发音模型 (AVAM) 直接对发音器官(例如嘴唇、舌头和牙齿)的运动进行建模。该模型采用具有共享发音器层的多流结构来同步关联语音的两个构建块,即音频和视频。该模型不仅描述了视觉发音运动和音频语音之间的同步性,而且反映了不同发音器官异步进化的语言事实。我们还提出了 Baum-Welch DBN 反演 (DBNI) 算法,根据最大似然 (ML) 标准下训练的 AVAM,根据音频生成最佳面部参数。对 JEWEL 视听数据集的广泛客观和主观评估表明,与音素 HMM 方法相比,我们的方法估计的面部参数更准确地遵循真实参数,并且合成的面部动画序列非常生动,其中 38% 无法区分
This paper presents an articulatory modelling approach to convert acoustic speech into realistic mouth animation. We directly model the movements of articulators, such as lips, tongue, and teeth, using a dynamic Bayesian network (DBN)-based audio-visual articulatory model (AVAM). A multiple-stream structure with a shared articulator layer is adopted in the model to synchronously associate the two building blocks of speech, i.e., audio and video. This model not only describes the synchronization between visual articulatory movements and audio speech, but also reflects the linguistic fact that different articulators evolve asynchronously. We also present a Baum-Welch DBN inversion (DBNI) algorithm to generate optimal facial parameters from audio given the trained AVAM under maximum likelihood (ML) criterion. Extensive objective and subjective evaluations on the JEWEL audio-visual dataset demonstrate that compared with phonemic HMM approaches, facial parameters estimated by our approach follow the true parameters more accurately, and the synthesized facial animation sequences are so lively that 38% of them are undistinguishable