课题基金 / 基金详情

HCC: Small: Modeling Acoustic and Articulatory Features for Hybrid Synthesis

HCC: Small: Modeling Acoustic and Articulatory Features for Hybrid Synthesis
HCC:小型:混合合成的声学和发音特征建模
批准号:
1116799
负责人:
Rupal Patel
金额:
$18.09万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-10-01 至 2014-09-30

项目摘要

项目成果

Rupal Patel的其他基金

相似基金

相关文献

中文摘要
翻译
近几十年来,合成语音已经成为人机界面的一个普遍存在且日益无缝的方面。 尽管汽车、微波炉、电话和信息亭都以类似人类的方式“说话”,但这些声音的自然性和个性却达不到人类的表达。 虽然这对于许多文本到语音(TTS)应用程序可能无关紧要,但超过200万患有严重语言运动障碍的美国人需要具有TTS输出的辅助通信辅助设备。 级联TTS合成器产生高度可理解的声音,但许多辅助设备依赖于小占地面积,共振峰合成,听起来像机器人,可理解性差。 此外,传统设备上的语音选择是有限的,并且不能反映用户;对于孩子来说,她一生都使用相同的语音,并且对于她的同龄人来说,即使在使用不同的设备时也共享相同的语音,这并不罕见。 这种缺乏关注的个性化合成语音的后果,通过辅助技术作为用户的延伸,并可能产生不利影响的社会态度对用户group.In她以前的工作PI开始解决这些问题,通过适应级联合成器构建从声学录音的健康说话者使用的声源特性获得的目标说话者与言语障碍。 适应的声音是高度可理解的,并传达了目标用户的身份,但它也保留了大量的元素,由于声道滤波器特性的影响,健康的谈话者的身份。 这表明,个性化的语音合成可能会更成功地利用替代方法,其中来自健康的谈话者的声学和发音数据与来自目标谈话者的源和滤波器特性相结合,以生成个性化的语音。 在该项目中,PI将开发混合统计参数合成技术,以模拟受损说话者的声道和源特征,目标是生成高度可理解和个性化的合成语音。 PI设想了一个未来,其中基于隐马尔可夫模型(HMM)的合成器的源和滤波器参数可以被适配为对儿童用户的声道进行建模,并随着时间的推移进行修改,以随着他成熟的声音系统“成长”,从而在用户和通信设备之间培养更强的个人联系。 该项目致力于通过设计一种模糊系统和用户之间界限的使能技术来实现通信的可访问性和社会满足感。 人类的声音不仅仅是一个信号;它具有个性化和个人的品质,影响他人如何看待我们以及我们如何与周围的人互动。 这项工作的最终目标是为TTS用户提供与自然语音相同的所有权和个性。 项目成果将对辅助器具使用者和TTS技术的健全使用者产生广泛影响。 这项研究也可能导致一种新的和创新的手段评估的性质和发音轨迹的语言障碍,通过比较模型参数受损的产品。 这项研究的跨学科性质将促进计算机科学和言语与听力科学的教学,培训和学习。
英文摘要
In recent decades synthetic speech has become a ubiquitous and increasingly seamless aspect of human-machine interfaces. Although cars, microwaves, phones, and kiosks all "talk" in human-like ways, the naturalness and personality of these voices fall short of human expression. While this may not matter for many text-to-speech (TTS) applications, over two million Americans with severe speech-motor impairments require assistive communication aids with TTS output. Concatenative TTS synthesizers yield highly intelligible voices, yet many assistive devices rely on small footprint, formant synthesis that sounds robotic and has poor intelligibility. Moreover, the choice of voices on conventional devices is limited and does not reflect the user; it is not uncommon for a child to use the same voice her whole life and for her peers to share that same voice even when using different devices. This lack of attention to the individuality of synthetic voices has consequences on adoption of assistive technology as an extension of the user, and may adversely impact societal attitudes toward the user group.In her prior work the PI began to address these issues by adapting a concatenative synthesizer constructed from acoustic recordings of a healthy talker using vocal source characteristics obtained from a target talker with speech impairment. The adapted voice was highly intelligible and conveyed the target user's identity, yet it also retained substantial elements of the healthy talker's identity due to the influence of vocal tract filter characteristics. This suggests that personalized speech synthesis may be more successful utilizing an alternative approach, in which acoustic and articulatory data from healthy talkers are combined with both source and filter characteristics from target talkers to generate an individualized voice. In this project, the PI will develop hybrid statistical parametric synthesis techniques to model vocal tract and source characteristics of impaired talkers, with the goal of generating highly intelligible and personalized synthetic speech. The PI envisages a future where source and filter parameters of a Hidden Markov Model (HMM) based synthesizer can be adapted to model a child user's vocal tract and modified over time to "grow" with his maturing vocal system, fostering a stronger personal connection between the user and the communication device.Broader Impacts: This project strives to make communication accessible and socially fulfilling by designing an enabling technology that blurs the line between system and user. The human voice is not merely a signal; it has an individualized and personal quality that impacts how others perceive us and how we interact with those around us. The ultimate goal of this work is to afford users of TTS the same ownership and individuality as the natural voice. Project outcomes will have broad impact both on users of assistive aids and able-bodied users of TTS technologies. The research may also lead to a novel and innovative means of assessing the nature and articulatory locus of speech impairment, by comparing model parameters to impaired productions. The interdisciplinary nature of this research will promote teaching, training and learning in computer science and in speech and hearing sciences.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SBIR Phase II: VocaliD - Infusing Unique Vocal Identities into Synthesized Speech
  • 批准号:
    1555608
  • 项目类别:
    Standard Grant
  • 资助金额:
    $74.78万
  • 财政年份:
    2016
  • 负责人:
    Rupal Patel
  • 依托单位:
SBIR Phase I: VocaliD - Infusing Unique Vocal Identities into Synthesized Speech
  • 批准号:
    1447995
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2015
  • 负责人:
    Rupal Patel
  • 依托单位:
EAGER: Collaborative Research: Wireless Sensing of Speech Kinematics and Acoustics for Remediation
  • 批准号:
    1449266
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2014
  • 负责人:
    Rupal Patel
  • 依托单位:
HCC-Small: Displaying Prosodic Text for Reading Aloud with Expression
  • 批准号:
    0915527
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $49.84万
  • 财政年份:
    2009
  • 负责人:
    Rupal Patel
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: