ITR-Collaborative Research: Development and Evaluation of a Hybrid Concatenative/Rule-Based Visual Speech Synthesis System
ITR-Collaborative Research: Development and Evaluation of a Hybrid Concatenative/Rule-Based Visual Speech Synthesis System
批准号:
0312434
负责人:
Lynne Bernstein
金额:
$21.68万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-07-15 至 2007-06-30
中文摘要
这个项目的目标是开发一个合成的会说话的脸。早在人类被要求阅读由计算机呈现的印刷文本之前,他们就发展出了感知和整合听觉和视觉(AV)语音信息的复杂能力。与只听说话者相比,看和听可以减少认知负荷,提高理解能力。为了实现AV语音在人机交互中的优势,需要合成视觉语音,从而在不需要预录制数据的情况下提供无限的视觉语音图像。这里的方法是用语音声学驱动光学语音合成。计算方法得到了从声学到光学的转换模型。该方法利用由听筒捕获的语音产生协同发音信息来产生自然的视觉语音图像。该方法直接应用于自然声语音特征,获得声光信号之间的协调关系。合成的视觉语音基于纹理映射线框模型。通过同时记录的三维光学、音频和视频数据,可以获得基于合成的自然语音语料库。综合开发以人类感知测试为指导。将分发DVD存档的文集。该项目将扩大获取信息的机会,并改善不同个人群体获取知识的情况,例如:仍在学习识字技能的儿童;识字能力不足的成年人;使用第二语言的个人;以及那些依赖视听语言的听力损失患者。结果将通过专业渠道广泛传播。研究生和本科生将参加。
英文摘要
This project's goal is to develop a synthetic talking face. Humans developed sophisticated abilities to perceive and integrate auditory and visual (AV) speech information long before they were required to read printed text presented by computers. Seeing as well as hearing speech reduces the cognitive workload and improves comprehension over only hearing the talker. To realize the advantages of AV speech for human-computer interactions requires synthesizing visual speech, thereby providing an unlimited supply of visual speech images without having to pre-record data. The approach here is to drive optical speech synthesis with speech acoustics. Computational methods obtain models of the transformation from acoustics to optics. The method capitalizes on the speech production coarticulatory information captured by diphones to produce naturalistic visual speech images. The method is applied directly to natural acoustic speech features to obtain coordination between acoustic and optical signals. The synthesized visual speech is based on a texture-mapped wire frame model. A natural speech corpus to base the synthesis is being obtained via simultaneously recorded 3-D optical, audio, and video data. Synthesis development is guided by human perceptual testing. The DVD archived corpus will be disseminated. The project will lead to expanded access to information and improvement in obtaining knowledge by diverse groups of individuals, for example: children still acquiring literacy skills; adults with inadequate literacy; individuals who are using a second language; and individuals with hearing losses who rely on audiovisual speech. Results will be disseminated broadly through professional outlets. Graduate and undergraduate students will participate.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
I-Corps: Smart Speech Perception Feedback for Training and Diagnostics
-
批准号:1738164
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2017
-
负责人:Lynne Bernstein
-
依托单位:
Collaborative Research: Using Somatosensory Speech And Non-Speech Categories To Test The Brain's General Principles Of Perceptual Learning
-
批准号:1439339
-
项目类别:Standard Grant
-
资助金额:$27.16万
-
财政年份:2014
-
负责人:Lynne Bernstein
-
依托单位:
Integration and Enhancement in Audiovisual Speech Perception
-
批准号:0214224
-
项目类别:Continuing Grant
-
资助金额:$34.61万
-
财政年份:2002
-
负责人:Lynne Bernstein
-
依托单位:
KDI: Segmental and Prosodic Optical Phonetics for Human and Machine Speech Processing
-
批准号:9872849
-
项目类别:Standard Grant
-
资助金额:$141.0万
-
财政年份:1998
-
负责人:Lynne Bernstein
-
依托单位:
海外基金