STTR Phase I: Small Footprint Speech Synthesis
STTR Phase I: Small Footprint Speech Synthesis
批准号:
0441125
负责人:
Alexander Kain
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2005
资助国家:
美国
项目状态:
已结题
起止时间:
2005-01-01 至 2006-06-30
中文摘要
这个小企业技术转移第一阶段项目旨在开发和实施文本到语音合成(TTS)领域的新算法,该算法将导致(I)在给定的语音质量水平下大幅降低磁盘和内存需求,以及(ii)将创建新的合成语音所需的录音量降至最低。大多数当前的TTS系统是通过连接录制的语音片段([声学]单元)来运行的。TTS的一个挑战是协同发音:一个音素的声学表现依赖于它的邻居。目前的TTS系统使用多电话声学单元,如双话筒,它保留了语音中自然存在的协同发音模式。然而,这种方法需要大量的记录,并且生成的系统占用空间很大。Biospeech提出了一种单一的方法,通过一个明确的模型来解决协同发音过程。该方法使用复杂的谱向量(基向量)表示单个音素内的简短语音片段,并将其分解为两个组成部分:形成峰向量和谱平衡向量。为了生成语音,从对应于连续音素的基向量派生的形成峰和谱平衡向量受到使用时变权重的单独(因此通常是异步的)插值操作;将由此产生的形成峰和谱平衡矢量轨迹重新组合在复谱空间中形成轨迹;最后,通过傅里叶反变换将该轨迹转换为输出语音。异步性是由不同频谱特征(如摩擦、共振峰频率)下发音器的准独立性所必需的。这项工作对其他语音技术也有启示,包括自动语音识别(ASR)。目前的ASR技术通过使用多电话单元(典型的三联电话)来解决协同发音问题。英语中三音琴的数量超过70,000,因此需要大量的培训录音。所提出的模型可能会极大地影响系统训练所需的录音量。第二,TTS已普遍认识到通过声音普及、教育和信息获取的社会效益。例如,基于tts的辅助设备可用于失声的个人;为盲人提供的阅读设备已经有几十年的历史了。第三,这种方法将使更高质量的TTS更适用于更小的设备。例如,由于内存限制,低端移动电话目前无法实现基于语音的来电显示。第四,它可以用最少的录音来适应语音。这将使建立个性化的TTS系统成为可能,适用于那些患有语言障碍的人,他们只能间歇性地发出正常的语音,或者是那些即将接受手术的人,这些手术将不可逆转地改变他们的语言。Biospeech提供的方法只需要记录每个(少于50个)音素的有效样本,而不是每个(2000个或更多)音素的有效样本。
英文摘要
This Small Business Technology Transfer Phase I project aims to develop and implement a new algorithm in the area of text-to-speech synthesis (TTS) that will lead to (i) dramatic decreases in disk and memory requirements at a given speech quality level and (ii) minimization of the amount of voice recordings needed to create a new synthetic voice. Most current TTS systems operate by concatenating segments of recorded speech ([acoustic] units). A challenge for TTS is coarticulation: The dependency of the acoustic manifestations of a phoneme on its neighbors. Current TTS systems use multi-phone acoustic units such as diphones, which preserve coarticulatory patterns naturally present in speech. However, this approach requires a large amount of recordings and generates systems with large footprints. Biospeech proposes a uniphone approach that addresses coarticulation processes with an explicit model. The method uses complex spectral vectors (basis vectors) representing brief segments of speech inside single phonemes, and decomposes these into two components: A formant vector and a spectral balance vector. To generate speech, the formant and spectral balance vectors derived from the basis vectors corresponding to successive phonemes are subjected to separate--and hence generally asynchronous--interpolation operations using time varying weights; the formant and spectral balance vector trajectories thus created are re-combined to create a trajectory in complex spectral space; finally, this trajectory is converted into output speech with the inverse Fourier transform. Asynchronicity is necessitated by the quasi-independence of articulators underlying different spectral features (e.g., frication, formant frequencies).The proposed work has implications for other speech technologies, including Automatic Speech Recognition (ASR). Current ASR technologies address coarticulation by using multi-phone units, typical triphones. The number of triphones in English is over 70,000, and thus requires a large amount of training recordings. The proposed model could dramatically impact on the amount of recordings required for system training. Second, TTS has generally recognized societal benefits for universal access, education, and information access by voice. For example, TTS-based augmentative devices are available for individuals who have lost their voice; and reading machines for the blind have been available for several decades. Third, the approach will make higher-quality TTS more available for smaller devices. For example, voice based caller ID on low-end mobile telephones is currently not possible due to memory limitations. Fourth, it enables voice adaptation with a minimum of recordings. This will enable building personalized TTS systems for individuals with speech disorders who can only intermittently produce normal speech sounds or for individuals who are about to undergo surgery that will irreversibly alter their speech. The method proffered by Biospeech only requires recordings of valid samples of each of (less than 50) phonemes instead of each of (2000 or more) diphones.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Medium: Collaborative Research: Semi-Supervised Discriminative Training of Language Models
-
批准号:0964102
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2010
-
负责人:Alexander Kain
-
依托单位:
Collaborative Research: CDI-Type I: Computational Models for the Automatic Recognition of Non-Human Primate Social Behaviors
-
批准号:1027834
-
项目类别:Standard Grant
-
资助金额:$57.78万
-
财政年份:2010
-
负责人:Alexander Kain
-
依托单位:
HCC: Medium: Synthesis and Perception of Speaker Identity
-
批准号:0964468
-
项目类别:Standard Grant
-
资助金额:$91.48万
-
财政年份:2010
-
负责人:Alexander Kain
-
依托单位:
RI: Small: Modeling Coarticulation for Automatic Speech Recognition
-
批准号:0915754
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2009
-
负责人:Alexander Kain
-
依托单位:
HCC: High-Quality Compression, Enhancement, and Personalization of Text-to-Speech Voices
-
批准号:0713617
-
项目类别:Continuing Grant
-
资助金额:$40.0万
-
财政年份:2007
-
负责人:Alexander Kain
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Baryogenesis, Dark Matter and Nanohertz Gravitational Waves from a Dark
Supercooled Phase Transition
-
批准号:24ZR1429700
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:YUICHIRO NAKAI
-
依托单位:
ATLAS实验探测器Phase 2升级
-
批准号:11961141014
-
项目类别:国际(地区)合作与交流项目
-
资助金额:3350万元
-
批准年份:2019
-
负责人:刘衍文
-
依托单位:
地幔含水相Phase E的温度压力稳定区域与晶体结构研究
-
批准号:41802035
-
项目类别:青年科学基金项目
-
资助金额:12.0万元
-
批准年份:2018
-
负责人:张里
-
依托单位:
基于数字增强干涉的Phase-OTDR高灵敏度定量测量技术研究
-
批准号:61675216
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2016
-
负责人:叶青
-
依托单位:
基于Phase-type分布的多状态系统可靠性模型研究
-
批准号:71501183
-
项目类别:青年科学基金项目
-
资助金额:17.4万元
-
批准年份:2015
-
负责人:陈童
-
依托单位:
纳米(I-Phase+α-Mg)准共晶的临界半固态形成条件及生长机制
-
批准号:51201142
-
项目类别:青年科学基金项目
-
资助金额:25.0万元
-
批准年份:2012
-
负责人:张英波
-
依托单位:
连续Phase-Type分布数据拟合方法及其应用研究
-
批准号:11101428
-
项目类别:青年科学基金项目
-
资助金额:23.0万元
-
批准年份:2011
-
负责人:黄卓
-
依托单位:
D-Phase准晶体的电子行为各向异性的研究
-
批准号:19374069
-
项目类别:面上项目
-
资助金额:6.4万元
-
批准年份:1993
-
负责人:张殿琳
-
依托单位: