课题基金 / 基金详情

Adapting a Text-to-Speech Synthesizer to Convey User Identity

Adapting a Text-to-Speech Synthesizer to Convey User Identity
采用文本转语音合成器来传达用户身份
批准号:
0712821
负责人:
Rupal Patel
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-07-15 至 2011-06-30

项目摘要

项目成果

Rupal Patel的其他基金

相似基金

相关文献

中文摘要
翻译
该项目将推进计算机化的语音合成方法,以便它们能够更好地近似单个说话人的独特声音特征。每个人的声音质量都是独一无二的,因此与个性、自我形象和他人的看法密不可分。我们可以改变自己的声音,让自己的声音听起来像其他人一样,吸引注意力,投射自信,传递权威,以及执行无数其他功能。即使使用最先进的文本到语音(TTS)合成,自然声音的这种灵活性也是不可想象的。虽然语音质量对于许多文本到语音的应用来说可能并不重要,但对于作为用户扩展的辅助通信辅助来说,语音质量是必不可少的。超过200万美国人患有严重的言语和运动障碍,需要他们使用辅助沟通辅助工具,其中许多人使用TTS来代表他们说话。市场上可用的设备上的语音输出选项极其有限。此外,合成语音在诸如年龄、性别、语音速率和语音质量等基本维度上不能代表用户,从而引起不必要的关注并削弱口头消息的内容,并且阻碍社会融合。这个项目的目的是利用严重言语障碍个人作品中的残余发声控制,以适应文本到语音合成器,以使合成的声音类似于用户的声音。考虑到用户语音损伤的严重性,传统的语音变形方法不能直接应用。最近的实证研究表明,患有严重言语障碍的儿童和成人保留了控制基本频率、口音、节奏和语速的能力,这些都是发出说话人身份信号的众多声学线索之一。这项研究将利用这种保留的能力来构建一种自适应的文本到语音合成器,在不降低可理解性的情况下传达用户的身份。带有身份的声音线索将从严重言语障碍的儿童中产生,并使用新的声音转换技术来适应年龄和性别匹配的串联合成声音。可用性测试将评估用户身份适应对TTS可理解性、自然度和可接受性的影响,其结果将为TTS适应的迭代设计提供见解。这项研究将对使用辅助工具的用户和使用语音合成的通信技术的健全用户产生更广泛的影响。该项目致力于通过设计一种使系统和用户之间的界限变得模糊的使人能够获得和满足社会的技术。最终目标是为语音合成技术的用户提供与自然声音相同的所有权和个性。这项工作的跨学科性质将促进计算机科学以及言语和听力科学的教学、培训和学习。
英文摘要
This project will advance computerized speech synthesis methods so that they can better approximate the unique vocal characteristics of individual human speakers. Voice quality is unique to each individual and thus is inextricable from personality, self-image, and the perceptions of others. We can alter our voice to sound like others, attract attention, project confidence, convey authority, and to perform countless other functions. This flexibility of the natural voice is inconceivable using even state-of-the art text-to-speech (TTS) synthesis. While voice quality may not matter for many text-to-speech applications, it is essential for assistive communication aids which are meant to be an extension of the user. Over two million Americans have severe speech and motor impairments that require them to use an assistive communication aid, many of whom use TTS to speak on their behalf. The speech output options on commercially available devices are extremely limited. Moreover, the synthetic voices are not representative of the user along basic dimensions such as age, gender, rate of speech, and voice quality thus drawing unnecessary attention and detracting from the content of the spoken message as well as impeding social integration. This project aims to harness the residual vocal control in the productions of individuals with severe speech impairment in order to adapt a text-to-speech synthesizer such that the resultant voice resembles that of the user. Conventional methods of voice morphing cannot be applied directly given the severity of the user's speech impairment. Recent empirical work suggests that children and adults with severe speech impairment retain the ability to control fundamental frequency, accent, rhythm, and speaking rate which are among the many acoustic cues that signal speaker identity. This research will leverage this preserved ability toward building an adaptive text-to-speech synthesizer that conveys the user's identity without degradation of intelligibility. Identity-bearing vocal cues will be elicited from children with severe speech impairment and used to adapt age and gender-matched concatenative synthetic voices using novel voice transformation techniques. Usability tests will be conducted to assess the impact of user identity adaptation on TTS intelligibility, naturalness, and acceptability, the results of which will provide insights for iterative design of TTS adaptation. The research will have broader impact on users of assistive aids and able-bodied users of communication technologies that use speech synthesis. This project strives to make communication accessible and socially fulfilling by designing an enabling technology in which the line between system and user is blurred. The ultimate goal is to afford users of speech synthesis technology the same ownership and individuality as the natural voice. The interdisciplinary nature of this work will promote teaching, training and learning in computer science and in speech and hearing sciences.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SBIR Phase II: VocaliD - Infusing Unique Vocal Identities into Synthesized Speech
  • 批准号:
    1555608
  • 项目类别:
    Standard Grant
  • 资助金额:
    $74.78万
  • 财政年份:
    2016
  • 负责人:
    Rupal Patel
  • 依托单位:
SBIR Phase I: VocaliD - Infusing Unique Vocal Identities into Synthesized Speech
  • 批准号:
    1447995
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2015
  • 负责人:
    Rupal Patel
  • 依托单位:
EAGER: Collaborative Research: Wireless Sensing of Speech Kinematics and Acoustics for Remediation
  • 批准号:
    1449266
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2014
  • 负责人:
    Rupal Patel
  • 依托单位:
HCC: Small: Modeling Acoustic and Articulatory Features for Hybrid Synthesis
  • 批准号:
    1116799
  • 项目类别:
    Standard Grant
  • 资助金额:
    $18.09万
  • 财政年份:
    2011
  • 负责人:
    Rupal Patel
  • 依托单位:
国内基金
海外基金
J-TEXT托卡马克上边界湍流与撕裂模相互作用的实验研究
  • 批准号:
    12375223
  • 项目类别:
    面上项目
  • 资助金额:
    54万元
  • 批准年份:
    2023
  • 负责人:
    刘海
  • 依托单位:
J-TEXT装置外加三维磁场主动调控偏滤器脱靶的实验研究
  • 批准号:
    12305243
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    20万元
  • 批准年份:
    2023
  • 负责人:
    周松
  • 依托单位:
J-TEXT托卡马克装置上多模式磁扰动对逃逸电流影响研究
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    林志芳
  • 依托单位:
J-TEXT托卡马克上边界湍流特性对高密度运行影响的实验研究
  • 批准号:
    11905080
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    26.0万元
  • 批准年份:
    2019
  • 负责人:
    石鹏
  • 依托单位: