HCC: High-Quality Compression, Enhancement, and Personalization of Text-to-Speech Voices
HCC: High-Quality Compression, Enhancement, and Personalization of Text-to-Speech Voices
批准号:
0713617
负责人:
Alexander Kain
金额:
$40.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-09-01 至 2011-08-31
中文摘要
人类语音信号的巨大可变性仍然是文本到语音(TTS)系统的核心挑战。本研究的目的是开发专注于消除拼接错误的TTS技术,并在协同发音、发音程度、韵律效果和说话人特征等方面进行准确的语音修改。研究人员正在探索一种异步插值模型(AIM),它有望提供高质量和灵活的TTS。AIM的核心思想是将一个短的语音区域表示为称为流的几种类型的特征的组合。每个流通过基向量的异步插值计算。每个基向量都与一个特定的音素、音素或更专门的单位相关联。因此,语音区域是通过几种类型的前后声学特征的不同程度的影响来描述的。利用AIM,研究人员还在开发方法,在给定大小或质量限制的情况下,优化压缩TTS系统的声学库存,并在给定一些训练样本的情况下使系统适应新的声音。所研究的系统将传统的串联合成与基于共振峰的合成相结合,将两者的优点结合在一起,形成了一个具有语音适应能力的高质量、优化的TTS系统。TTS在通过声音获得教育和信息等方面的社会效益已得到普遍认可。例如,我们的研究将使为那些只能间歇性发出正常语音的言语障碍患者建立个性化的TTS系统成为可能。
英文摘要
The vast variability of the human speech signal remains a central challenge for Text-to-Speech (TTS) systems. The objective of this research is to develop TTS technologies that focus on elimination of concatenation errors, and accurate speech modifications in the areas of coarticulation, degree of articulation, prosodic effects, and speaker characteristics. The investigators are exploring an asynchronous interpolation model (AIM), which promises to provide for high-quality and flexible TTS. The core idea of AIM is to represent a short region of speech as a composition of several types of features called streams.Each stream is computed by asynchronous interpolation of basis vectors.Each basis vector is associated with a particular phoneme, allophone, or more specialized unit. Thus, the speech region is described by the varying degrees of influence of several types of preceding and following acoustic features. Using AIM, the investigators are also developing methods to optimally compress the acoustic inventories of TTS systems, given a size or a quality constraint, and to adapt the system to a new voice, given a few training samples. The system being researched forms a hybrid between traditional concatenative and formant-based synthesis, having advantages of both, resulting in a high-quality, optimized TTS system with voice adaptation capabilities. TTS has generally recognized societal benefits for universal access, education, and information access by voice. Our research will make it possible, for example, to build personalized TTS systems for individuals with speech disorders who can only intermittently produce normal speech sounds.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Medium: Collaborative Research: Semi-Supervised Discriminative Training of Language Models
-
批准号:0964102
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2010
-
负责人:Alexander Kain
-
依托单位:
Collaborative Research: CDI-Type I: Computational Models for the Automatic Recognition of Non-Human Primate Social Behaviors
-
批准号:1027834
-
项目类别:Standard Grant
-
资助金额:$57.78万
-
财政年份:2010
-
负责人:Alexander Kain
-
依托单位:
HCC: Medium: Synthesis and Perception of Speaker Identity
-
批准号:0964468
-
项目类别:Standard Grant
-
资助金额:$91.48万
-
财政年份:2010
-
负责人:Alexander Kain
-
依托单位:
RI: Small: Modeling Coarticulation for Automatic Speech Recognition
-
批准号:0915754
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2009
-
负责人:Alexander Kain
-
依托单位:
STTR Phase I: Small Footprint Speech Synthesis
-
批准号:0441125
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Alexander Kain
-
依托单位:
海外基金