Combined X-ray and facial videos for phoneme-level articulator dynamics

Combined X-ray and facial videos for phoneme-level articulator dynamics
复制标题

DOI:
10.1007/s00371-010-0434-1
复制
发表时间:
2010-06
期刊:
The Visual Computer
影响因子:
--
通讯作者:
Hui Chen;Lan Wang;Wenxi Liu;P. Heng
Hui Chen;Lan Wang;Wenxi Liu;P. Heng
中科院分区:
其他
文献类型:
--
作者:
Hui Chen;Lan Wang;Wenxi Liu;P. Heng

文献摘要

被引文献

相似文献

动态外部和内部咬合架运动集成到一个低成本的数据驱动的三维说话头在本文中。外部和内部的关节定义和校准的视频流和视频透视到一个通用的3D说话的头部模型。根据唇舌肌软组织、下巴上下运动和相对固定的发音器官的发音特点,建立并整合了三种不同的变形模式。将自然语音输入的分段音素之间的形状混合函数合成到一个话语中。易混淆音素和最小音对的动画展示给英语教师和学习者进行感知测试。结果表明,该方法能够真实地反映语音发音的真实的情况。
Dynamic external and internal articulator motions are integrated into a low-cost data-driven three-dimensional talking head in this paper. External and internal articulations are defined and calibrated from the video streams and the videofluoroscopy to a generic 3D talking head model. Three different deformation modes in relation to pronunciation characteristics of muscular soft tissue of lips and tongue, up-down movements of chin and the relatively fixed articulators are set up and integrated. The shape blending functions among segmented phonemes of natural speech input are synthesized in an utterance. Animations of the confusable phonemes and minimal pairs are shown to English teachers and learners for a perception test. The results show that the proposed method can reflect the real situation of phonetic pronunciation realistically.