TRANSFORM: flexible voice synthesis through articulatory voice transformation
TRANSFORM: flexible voice synthesis through articulatory voice transformation
批准号:
0414675
负责人:
Alan Black
金额:
$0.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-05-15 至 2009-04-30
中文摘要
许多人一直希望机器能与他们交谈,但大多数人对特定的声音有强烈的偏好。 目前的语音合成技术可以构建听起来非常接近原始说话者的声音,捕捉源声音的风格,方式和清晰度。 然而,这样的系统需要许多小时的仔细记录的语音和专家调音,以达到可接受的质量水平。 一个令人兴奋的新的替代方法来构建合成语音是语音转换。 这种方法使用现有的录音数据库,并使用10-20个句子将其转换为目标语音。 这项技术提供了使语音合成器以任何所需的声音说话的潜力,比以前的技术所需的努力要少得多。 当前的变换技术集中在语音的频谱映射上,即,转换所述语音信号的属性。 相反,我们使用声道发音器官的潜在位置(即,牙齿、舌头、嘴唇、软腭的位置),这引起语音的频谱输出。 使用新的统计建模技术,我们可以成功地预测的位置发言人的发音从语音信号。 然后在虚拟声道域中映射说话人之间的语音,并为目标语音重新生成语音。这项工作使得能够轻松构建新的合成语音,从而实现语音输出的个性化。 它增加了我们对语音生成过程的了解,并描述了使语音个性化的特征。
英文摘要
Many people have always wanted machines to talk to them, but most have strong preferences for particular voices. Current techniques in speech synthesis can build voices that sound very close to the original speaker, capturing the style, manner and articulation of the source voice. However such systems require many hours of carefully recorded speech and expert tuning to reach an acceptable level of quality. An exciting new alternative method for building synthetic voices is voice transformation. This method uses an existing recorded database and converts it to a target voice using as little as 10-20 sentences. This technique offers the potential to make speech synthesizers talk in whatever voice desired, with significantly less effort required than previous techniques.This project offers a new direction in voice transformation. Current transformation techniques concentrate on a spectral mapping of the voice, i.e., converting the properties of the speech signal. Instead we use the underlying positions of the vocal tract articulators (i.e., the position of the teeth, tongue, lips, velum), which give rise to the spectral output of the voice. Using new statistical modeling techniques we can successfully predict the positions of a speaker's articulators from the speech signal. Then in the virtual vocal tract domain map between speakers and regenerate the speech for the target voice.This work enables the easy construction of new synthetic voices allowing personalization of speech output. It increases our knowledge of the speech generation process and characterizes what make a voice personal.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Modeling Lexical Borrowing to Bridge the "Linguistic Divide" in Natural Language Processing
-
批准号:1526745
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2015
-
负责人:Alan Black
-
依托单位:
ITR: Evaluation and Personalization of Synthetic Voices
-
批准号:0219687
-
项目类别:Continuing Grant
-
资助金额:$39.43万
-
财政年份:2002
-
负责人:Alan Black
-
依托单位:
国内基金
海外基金
A study on prototype flexible multifunctional graphene foam-based sensing grid (柔性多功能石墨烯泡沫传感网格原型研究)
-
批准号:--
-
项目类别:--
-
资助金额:20万元
-
批准年份:2020
-
负责人:SAGAR RIZWAN UR REHMAN
-
依托单位: