课题基金 / 基金详情

项目摘要

项目成果

SUSAN R HERTZ的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):NovaSpeech提出开发一种创新的以感知为导向的混合方法,用于无约束语音合成,以生成个性化,定制的任何性别和任何年龄的声音。该系统将提供听起来像人的、可理解的和模仿的语音,但存储需求很小,能够支持低成本的新语音添加,并且适合在几乎任何硬件平台上实现。因此,这项技术将非常适合几乎任何无限的词汇合成应用,但对语言障碍人士尤其有利,因为他们特别需要在各种各样的设备上听到听起来自然、个性化的声音。有了这个混合系统,那些知道自己将因疾病或手术而失声的人将能够经济高效地在语音输出通信辅助设备中捕捉和利用他们受伤前的声音;所有有语言障碍的用户都将能够获得可靠的、适当的、个性化的声音,随着他们的成熟和年龄的增长,这些声音可以与他们一起成长。现有的合成方法都无法满足这些需求,每种技术都在用一种可取的特性换取另一种可取的特性,无论是自然语音质量的低存储要求,还是灵活性的人类语音质量。混合方法克服了这些限制,以一种新颖和原则性的方式,整合了两种知名合成技术的最佳特征:基于语料库的波形拼接和基于规则的形成峰合成。利用一些重要的感知原则,该系统将只预先存储来自目标说话者的少量固有单元,如重读元音,并根据规则合成其他可适应的单元。因此,只需要一个小的预先存储的语音语料库和一套通用的语音规则,它就会产生听起来像预期说话者的语音。在其拟议的第二阶段项目中,NovaSpeech将开发一个完整的混合原型文本到语音(TTS)系统,用于通用美式英语中的八种声音,包括男性和女性儿童、成人和老年人(基本说话者),以及两名知道自己将因喉切除术而失去自然说话能力的说话者。第一年将专注于探索可能的系统架构;适应性单位实施细则;并通过知觉实验探索存储和选择内在单位的可能策略。第二年将专注于实现六个基本声音的全功能混合TTS原型。最迟在第二年的第六个月,公司将通过实施喉切除术患者的声音,为他们提供声音的功能系统,并从他们和了解他们的人那里获得声音质量和系统功能的反馈,来验证快速添加新声音的能力。该混合项目的最终目标是提高由无限制符号输入合成的语音的自然度和模仿质量,特别是提高语音输出通信辅助设备的实用性和灵活性。
英文摘要
DESCRIPTION (provided by applicant): NovaSpeech proposes to develop an innovative perceptually-oriented hybrid approach to unconstrained speech synthesis for generating individualized, customized voices of either gender and any age. The system will provide human-sounding, intelligible, and mimetic speech, yet have small storage requirements, be able to support the cost-efficient addition of new voices, and be suitable for implementation on virtually any hardware platform. As a result, the technology will be well-suited to virtually any unlimited vocabulary synthesis application, but be of special benefit to speech-impaired individuals, who have a particularly great need for natural-sounding, individualized voices on a broad range of devices. With the hybrid system, individuals who know they will lose their voice due to illness or surgery will be able to cost-efficiently capture and utilize their pre-injury voice in a voice output communication aid; and all speech-impaired users will be able to obtain reliable, appropriate, individualized voices that can grow with them as they mature and age. No existing synthesis approach meets these needs, with each type of technology trading off one desirable property for another, be it low storage requirements for natural voice quality, or human voice quality for flexibility. The hybrid approach overcomes these limitations by integrating, in a novel and principled way, the best features of two well-known synthesis techniques: corpus-based waveform concatenation and rule-based formant synthesis. Capitalizing on a number of important perceptual principles, the system will prestore only a small number of intrinsic units, such as stressed vowels, from the target speaker, and synthesize other, adaptable units by rule. Thus with only a small prestored speech corpus, and a common set of rules across voices, it will produce speech that sounds like the intended speaker. In its proposed Phase II project, NovaSpeech will develop a complete hybrid prototype text-to-speech (TTS) system for eight voices in General American English, including male and female children, adults, and elderly adults (the base speakers), as well as for two speakers who know they will lose their ability to speak naturally as a result of future laryngectomies. Year 1 will be focused on exploring possible system architectures; implementing rules for adaptable units; and exploring through perceptual experiments possible strategies for storing and selecting intrinsic units. Year 2 will be focused on implementing a fully functional hybrid TTS prototype for the six base voices. By month six of year 2 at the latest, the company will verify the ability to quickly add new voices by implementing the voices of the laryngectomy patients, providing them with functional systems for their voices, and obtaining feedback from them and those who know them about the quality of the voices and system features. The ultimate objective of the hybrid project is to improve the naturalness and mimetic quality of speech synthesized from unrestricted symbolic input, with the particular goal of enhancing the utility and flexibility of voice output communication aids for speech-impaired individuals.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Expressive Speech Synthesis for Speech-Generating Devices
  • 批准号:
    8903390
  • 项目类别:
  • 资助金额:
    $15.0万
  • 财政年份:
    2015
  • 负责人:
    SUSAN R HERTZ
  • 依托单位:
Hybrid Synthesis For Voice Output Communication Aids
  • 批准号:
    6790229
  • 项目类别:
  • 资助金额:
    $10.0万
  • 财政年份:
    2004
  • 负责人:
    SUSAN R HERTZ
  • 依托单位:
Hybrid Speech Synthesis for Voice Output Communication Aids
  • 批准号:
    7156322
  • 项目类别:
  • 资助金额:
    $37.35万
  • 财政年份:
    2004
  • 负责人:
    SUSAN R HERTZ
  • 依托单位:
OPTIMIZATION OF SPEECH SYNTHESIS SOFTWARE
  • 批准号:
    2252069
  • 项目类别:
  • 资助金额:
    $25.65万
  • 财政年份:
    1991
  • 负责人:
    SUSAN R HERTZ
  • 依托单位:
海外基金