Next-Generation Expressive Personalized Voices for Speech-Generating Devices
Next-Generation Expressive Personalized Voices for Speech-Generating Devices
批准号:
10547241
负责人:
H TIMOTHY Bunnell
金额:
$27.58万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
已结题
起止时间:
2022-08-15 至 2024-08-14
关键词:
ALS patientsAdoptionAdultAgeAlgorithmsAmyotrophic Lateral SclerosisAugmentative and Alternative CommunicationCharacteristicsChildChild HealthClientDepressed moodDiseaseDysarthriaEmotionsEncapsulatedEvaluationFemaleGenerationsGoalsGovernmentHumanHybridsIndividualKnowledgeLaboratory ResearchLearningLinguisticsMachine LearningMethodsModelingNetwork-basedNeurodegenerative DisordersOnset of illnessOutcomeOutputPersonsPhaseProcessProductionReadingRecordsRehabilitation therapyRiskRunningServicesSpeechStructureSurveysSystemTechnologyTextTrainingVoiceVoice Qualitybasecommercial applicationcommunication devicedeep neural networkdesignexperienceexperimental studyimprovedknowledge basemachine learning algorithmmalemimeticsnext generationnovelsoundsuccessvirtual vocal tract
中文摘要
项目摘要/摘要
个性化合成语音的创建在医疗/康复环境中具有广泛的应用。
依靠语音生成设备(SGD)进行交流的人。一种常见的应用是语音
银行业,指有失声风险的人,例如患有神经退行性疾病的人
像肌萎缩侧索硬化症(ALS)一样,在疾病相关的Dysar发作之前记录自己的语言-
用于以后在模拟其自然语音特征的SGD中使用的Thria。虽然潜在的技术
这种个性化的合成声音的创作越来越成熟,并被SGD用户采用,但它仍然适用于
FERS来自两个主要限制:缺乏表现力和需要繁重的记录量
创造出非常自然的声音。拟议的项目旨在通过与马-马结婚来纠正这种情况。
ModelTalker背后的中文学习技术,ModelTalker是一种开创性的语音银行文本到语音转换服务,在
Nemour儿童健康,基于知识的技术支持Synfony,一种基于规则的文本到
由Synfonica LLC开发的语音系统,它能够生成各种语音风格和前缀。
压倒性的模式。Synfonica内置的专家知识将被用于设计一组最佳的句子
供语音银行家录制,其算法用于生成听起来自然的不同韵律
模式和风格将被集成到ModelTalker的机器学习算法中,创建一个混合系统
这包含了这两种方法的最佳品质。由此产生的新的文本到语音(TTS)系统
项目将(A)需要来自语音管理员的最少量的录音语音,(B)准确捕获
他们的声音身份,以及(C)结构,以便可以容易地添加新的表达模式和讲话风格
不需要额外的录制。该项目的可行性将通过录制一位
成年男性、成年女性和儿童,并产生可以三种表达模式说话的TTS声音
(中性、快乐和悲伤)。将进行感知实验,以评估它们的可理解性、自然度、舒适性-
成功地捕捉说话人的声音身份,以及他们表达方式的适当性。在这一代-
此外,该项目将是使个性化合成语音的用户能够表达
他们的情绪和意图。
英文摘要
Project Summary/Abstract
The creation of personalized synthetic voices has wide application in medical/rehabilitation settings for pa-
tients who rely on a speech-generating device (SGD) for communication. One common application is voice
banking, wherein a person who risks losing their voice, such as somebody with a neurodegenerative disease
like Amyotrophic Lateral Sclerosis (ALS), records their own speech before the onset of disease-related dysar-
thria for later use in an SGD that mimics their natural speech characteristics. While the technology underlying
the creation of such personalized synthetic voices is growing in maturity and adoption by SGD users, it still suf-
fers from two primary limitations: a lack of expressiveness and a burdensome amount of recording needed to
create highly natural-sounding voices. The proposed project aims to remedy this situation by marrying the ma-
chine-learning technology behind ModelTalker, a pioneering voice-banking text-to-speech service developed at
Nemours Children’s Health, with the knowledge-based technology underlying Synfony, a rule-based text-to-
speech system developed by Synfonica LLC, which is capable of generating a variety of speech styles and ex-
pressive modes. The expert knowledge built into Synfonica will be used to design an optimal set of sentences
for voice bankers to record, and its algorithms for the generation of natural-sounding prosody in different
modes and styles will be integrated into ModelTalker’s machine-learning algorithms, creating a hybrid system
that embraces the best qualities of both approaches. The new text-to-speech (TTS) system resulting from this
project will (a) require a minimal amount of recorded speech from the voice banker, (b) accurately capture
their vocal identity, and (c) be structured such that new expressive modes and speech styles can be added easily
without additional recording. The feasibility of the project will be demonstrated by recording the voices of an
adult male, an adult female, and a child, and generating TTS voices that can speak in three expressive modes
(neutral, happy, and sad). Perceptual experiments will be run to evaluate their intelligibility, naturalness, suc-
cess in capturing the vocal identity of the speaker, and the appropriateness of their expressive modes. In gen-
eral, the project will be a major step forward in enabling the users of personalized synthetic voices to express
their emotions and intentions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Phenotypes and familiality in speech disorders
-
批准号:6908541
-
项目类别:
-
资助金额:$15.94万
-
财政年份:2005
-
负责人:H TIMOTHY Bunnell
-
依托单位:
Phenotypes and familiality in speech disorders
-
批准号:7038247
-
项目类别:
-
资助金额:$15.56万
-
财政年份:2005
-
负责人:H TIMOTHY Bunnell
-
依托单位:
Personalized speech output for communication devices
-
批准号:7219783
-
项目类别:
-
资助金额:$41.79万
-
财政年份:2003
-
负责人:H TIMOTHY Bunnell
-
依托单位:
Personalizing Speech Output for Communication Devices
-
批准号:6749031
-
项目类别:
-
资助金额:$9.99万
-
财政年份:2003
-
负责人:H TIMOTHY Bunnell
-
依托单位:
Personalizing Speech Output for Communication Devices
-
批准号:6646704
-
项目类别:
-
资助金额:$9.99万
-
财政年份:2003
-
负责人:H TIMOTHY Bunnell
-
依托单位:
Personalized speech output for communication devices
-
批准号:7337320
-
项目类别:
-
资助金额:$42.15万
-
财政年份:2003
-
负责人:H TIMOTHY Bunnell
-
依托单位:
海外基金