Adapting a Text-to-Speech Synthesizer to Convey User Identity
Adapting a Text-to-Speech Synthesizer to Convey User Identity
批准号:
0712821
负责人:
Rupal Patel
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-07-15 至 2011-06-30
中文摘要
该项目将推进计算机语音合成方法,以便更好地近似个人说话者的独特声音特征。 声音质量对每个人来说都是独一无二的,因此与个性,自我形象和他人的看法密不可分。我们可以改变自己的声音听起来像别人,吸引注意力,项目的信心,传达权威,并执行无数其他功能。 自然语音的这种灵活性即使使用最先进的文本到语音(TTS)合成也是不可想象的。虽然语音质量对于许多文本到语音的应用程序来说可能并不重要,但对于旨在成为用户扩展的辅助通信辅助设备来说,它是必不可少的。超过200万美国人有严重的语言和运动障碍,需要使用辅助沟通工具,其中许多人使用TTS来代表他们说话。商用设备上的语音输出选项极其有限。此外,合成语音并不代表用户的沿着基本维度,例如年龄、性别、语速和语音质量,因此引起不必要的注意,并从口头消息的内容中转移注意力,以及阻碍社会融合。该项目旨在利用剩余的语音控制在生产的个人严重的语言障碍,以适应文本到语音合成器,使合成的声音类似于用户。传统的语音变形方法不能直接应用于给定的用户的言语障碍的严重性。最近的实证研究表明,严重言语障碍的儿童和成人保留了控制基频、口音、节奏和语速的能力,这些都是发出说话者身份信号的声学线索。这项研究将利用这种保留的能力,建立一个自适应的文本到语音合成器,传达用户的身份,而不降低可理解性。身份轴承的声音线索将引起严重的言语障碍的儿童,并用于适应年龄和性别匹配的拼接合成声音,使用新的语音转换技术。可用性测试将进行评估的影响,用户身份的TTS的可懂度,自然度和可接受性,其结果将提供见解的TTS适应迭代设计。 这项研究将对使用辅助工具的用户和使用语音合成的通信技术的健全用户产生更广泛的影响。该项目致力于通过设计一种使系统和用户之间的界限模糊的技术,使通信变得容易获得和社会满意。最终目标是为语音合成技术的用户提供与自然声音相同的所有权和个性。这项工作的跨学科性质将促进计算机科学和言语与听力科学的教学、培训和学习。
英文摘要
This project will advance computerized speech synthesis methods so that they can better approximate the unique vocal characteristics of individual human speakers. Voice quality is unique to each individual and thus is inextricable from personality, self-image, and the perceptions of others. We can alter our voice to sound like others, attract attention, project confidence, convey authority, and to perform countless other functions. This flexibility of the natural voice is inconceivable using even state-of-the art text-to-speech (TTS) synthesis. While voice quality may not matter for many text-to-speech applications, it is essential for assistive communication aids which are meant to be an extension of the user. Over two million Americans have severe speech and motor impairments that require them to use an assistive communication aid, many of whom use TTS to speak on their behalf. The speech output options on commercially available devices are extremely limited. Moreover, the synthetic voices are not representative of the user along basic dimensions such as age, gender, rate of speech, and voice quality thus drawing unnecessary attention and detracting from the content of the spoken message as well as impeding social integration. This project aims to harness the residual vocal control in the productions of individuals with severe speech impairment in order to adapt a text-to-speech synthesizer such that the resultant voice resembles that of the user. Conventional methods of voice morphing cannot be applied directly given the severity of the user's speech impairment. Recent empirical work suggests that children and adults with severe speech impairment retain the ability to control fundamental frequency, accent, rhythm, and speaking rate which are among the many acoustic cues that signal speaker identity. This research will leverage this preserved ability toward building an adaptive text-to-speech synthesizer that conveys the user's identity without degradation of intelligibility. Identity-bearing vocal cues will be elicited from children with severe speech impairment and used to adapt age and gender-matched concatenative synthetic voices using novel voice transformation techniques. Usability tests will be conducted to assess the impact of user identity adaptation on TTS intelligibility, naturalness, and acceptability, the results of which will provide insights for iterative design of TTS adaptation. The research will have broader impact on users of assistive aids and able-bodied users of communication technologies that use speech synthesis. This project strives to make communication accessible and socially fulfilling by designing an enabling technology in which the line between system and user is blurred. The ultimate goal is to afford users of speech synthesis technology the same ownership and individuality as the natural voice. The interdisciplinary nature of this work will promote teaching, training and learning in computer science and in speech and hearing sciences.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SBIR Phase II: VocaliD - Infusing Unique Vocal Identities into Synthesized Speech
-
批准号:1555608
-
项目类别:Standard Grant
-
资助金额:$74.78万
-
财政年份:2016
-
负责人:Rupal Patel
-
依托单位:
SBIR Phase I: VocaliD - Infusing Unique Vocal Identities into Synthesized Speech
-
批准号:1447995
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2015
-
负责人:Rupal Patel
-
依托单位:
EAGER: Collaborative Research: Wireless Sensing of Speech Kinematics and Acoustics for Remediation
-
批准号:1449266
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2014
-
负责人:Rupal Patel
-
依托单位:
HCC: Small: Modeling Acoustic and Articulatory Features for Hybrid Synthesis
-
批准号:1116799
-
项目类别:Standard Grant
-
资助金额:$18.09万
-
财政年份:2011
-
负责人:Rupal Patel
-
依托单位:
HCC-Small: Displaying Prosodic Text for Reading Aloud with Expression
-
批准号:0915527
-
项目类别:Continuing Grant
-
资助金额:$49.84万
-
财政年份:2009
-
负责人:Rupal Patel
-
依托单位:
SGER: Loudmouth - Toward Intelligible Speech Synthesis in Everyday Noise Situations
-
批准号:0509935
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Rupal Patel
-
依托单位:
国内基金
海外基金
登录
查看更多内容
J-TEXT托卡马克上边界湍流与撕裂模相互作用的实验研究
-
批准号:12375223
-
项目类别:面上项目
-
资助金额:54万元
-
批准年份:2023
-
负责人:刘海
-
依托单位:
J-TEXT装置外加三维磁场主动调控偏滤器脱靶的实验研究
-
批准号:12305243
-
项目类别:青年科学基金项目
-
资助金额:20万元
-
批准年份:2023
-
负责人:周松
-
依托单位:
J-TEXT托卡马克装置上多模式磁扰动对逃逸电流影响研究
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:林志芳
-
依托单位:
J-TEXT托卡马克上边界湍流特性对高密度运行影响的实验研究
-
批准号:11905080
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2019
-
负责人:石鹏
-
依托单位:
关于J-TEXT托卡马克上微撕裂模电磁湍流及其输运的实验研究
-
批准号:11605067
-
项目类别:青年科学基金项目
-
资助金额:19.0万元
-
批准年份:2016
-
负责人:陈杰
-
依托单位:
基于J-TEXT远红外偏振干涉仪的相干散射与密度扰动的实验研究
-
批准号:11575067
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:高丽
-
依托单位:
J-TEXT上外加磁扰动抑制等离子体破裂下逃逸电子产生的实验研究
-
批准号:11275079
-
项目类别:面上项目
-
资助金额:80.0万元
-
批准年份:2012
-
负责人:陈忠勇
-
依托单位:
J-TEXT托卡马克等离子体粒子输运的密度调制实验研究
-
批准号:11105056
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2011
-
负责人:高丽
-
依托单位: