课题基金 / 基金详情

HCC: High-Quality Compression, Enhancement, and Personalization of Text-to-Speech Voices

HCC: High-Quality Compression, Enhancement, and Personalization of Text-to-Speech Voices
HCC:文本转语音的高质量压缩、增强和个性化
批准号:
0713617
负责人:
Alexander Kain
金额:
$40.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-09-01 至 2011-08-31

项目摘要

项目成果

Alexander Kain的其他基金

相似基金

相关文献

中文摘要
翻译
人类语音信号的巨大可变性仍然是文本到语音(TTS)系统的核心挑战。本研究的目的是开发TTS技术,专注于消除拼接错误,准确的语音修改的协同发音,清晰度,韵律效果和扬声器特性的领域。研究人员正在探索一种异步插值模型(AIM),该模型有望提供高质量和灵活的TTS。AIM的核心思想是将语音的一个短区域表示为几种称为流的特征的组合。每个流通过基向量的异步插值来计算。每个基向量与特定的音素、音位变体或更专门的单元相关联。因此,语音区域是由几种类型的前后声学特征的不同程度的影响来描述的。使用AIM,研究人员还在开发方法,以最佳地压缩TTS系统的声学库存,给定大小或质量约束,并使系统适应新的声音,给定一些训练样本。该系统融合了传统拼接合成和基于共振峰合成的优点,实现了具有语音自适应能力的高质量、优化的TTS系统。TTS普遍认识到通过语音实现普遍接入、教育和信息接入的社会效益。例如,我们的研究将使为那些只能间歇性地发出正常语音的言语障碍患者建立个性化的TTS系统成为可能。
英文摘要
The vast variability of the human speech signal remains a central challenge for Text-to-Speech (TTS) systems. The objective of this research is to develop TTS technologies that focus on elimination of concatenation errors, and accurate speech modifications in the areas of coarticulation, degree of articulation, prosodic effects, and speaker characteristics. The investigators are exploring an asynchronous interpolation model (AIM), which promises to provide for high-quality and flexible TTS. The core idea of AIM is to represent a short region of speech as a composition of several types of features called streams.Each stream is computed by asynchronous interpolation of basis vectors.Each basis vector is associated with a particular phoneme, allophone, or more specialized unit. Thus, the speech region is described by the varying degrees of influence of several types of preceding and following acoustic features. Using AIM, the investigators are also developing methods to optimally compress the acoustic inventories of TTS systems, given a size or a quality constraint, and to adapt the system to a new voice, given a few training samples. The system being researched forms a hybrid between traditional concatenative and formant-based synthesis, having advantages of both, resulting in a high-quality, optimized TTS system with voice adaptation capabilities. TTS has generally recognized societal benefits for universal access, education, and information access by voice. Our research will make it possible, for example, to build personalized TTS systems for individuals with speech disorders who can only intermittently produce normal speech sounds.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Medium: Collaborative Research: Semi-Supervised Discriminative Training of Language Models
Collaborative Research: CDI-Type I: Computational Models for the Automatic Recognition of Non-Human Primate Social Behaviors
HCC: Medium: Synthesis and Perception of Speaker Identity
RI: Small: Modeling Coarticulation for Automatic Speech Recognition
海外基金