课题基金 / 基金详情

CHS: Small: Compounding Dividends on Voice Banking

CHS: Small: Compounding Dividends on Voice Banking
CHS:小:语音银行的复利红利
批准号:
1816726
负责人:
H. Timothy Bunnell
金额:
$10.41万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-03-01 至 2022-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
文语转换(TTS)合成已经成为一种成功的、普遍存在的技术。激励本研究的TTS技术的应用领域是其用于增强和替代通信(AAC)。根据美国语言和听力协会(阿莎)的数据,美国有超过200万人患有严重的沟通障碍,这损害了他们的说话能力。使用TTS来创建语音输出的AAC设备被这些人中的许多人用来支持通信。从历史上看,AAC用户可以访问相对较小的通用TTS语音家族,这些语音既不是他们独有的,也不是典型的年龄或方言。然而,TTS技术的进步使得创建个性化的合成语音成为可能,如果他们能够录制足够的语音,则可以捕获AAC设备用户的独特声音身份。这使得患有神经退行性疾病(如ALS)的患者能够“储存”他们的声音-也就是说,在疾病发展到无法说话之前,记录他们的语音示例,以便稍后用于创建个人TTS语音。不幸的是,语音银行的一个主要障碍,特别是对于那些可能已经经历了一些说话困难的患者,是创建一个自然的TTS声音所需的语音量,完全捕捉语音银行的声音身份。为了减少这一障碍,这项研究将联合收割机结合一种称为并行共振峰合成的语音合成技术,这种技术是几十年前开发的,深度学习计算技术允许计算机学习如何控制并行共振峰合成器的参数,以重现目标说话者的语音。一个并行共振峰合成器将被实现和训练,以模拟语音银行记录的语音,其输出将与其他合成器,已与相同的语音数据进行了训练。将使用合成和自然话语之间的相似性的客观度量,以及使用人类听众的语音质量和相似性的主观度量。这将是建立一个并行共振峰合成为基础的语音转换系统的第一步,能够创建TTS语音从少量的自然语音样本,也能够更好地模拟自然speech.Despite TTS技术的进步,语音银行的应用这项技术有多种挑战。具体而言:(a)所需的发言量(B)与级联合成相比,不需要来自目标说话者的大量并行语音的现有语音转换技术通常产生听起来不太自然且不太像目标说话者的语音;以及(c)连接和统计参数技术都产生仅与语音语料库内的数据一样有表现力的语音,其中从所述语音语料库构建或训练所述连接和统计参数技术。并行共振峰合成,因为它是基于自然语音的感知最显着的特征,并适合于独立建模喉,超音段,和分段功能,应该能够更好地解决所有这三个挑战。作为概念证明,将实现具有基于DNN的参数估计的并行共振峰合成(PFS)声码器。声码器将在Merlin DNN合成框架内实现,以便PFS系统的语音输出可以直接与World和MagPhase声码器生成的输出进行比较。培训将基于语料库,这些语料库来自多个人记录的1600个话语,这些人将他们的录音贡献给ModelTalker项目。选定的目标说话者将在性别上保持平衡,并涵盖广泛的英语方言,但将避免使用具有明显构音障碍水平的说话者。客观的比较将基于在训练合成器时未使用的合成和自然句子标记之间的梅尔倒谱差(MCD)。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Text to speech (TTS) synthesis has become a successful and ubiquitous technology. The area of application for TTS technology that motivates this research is its use for Augmentative and Alternative Communication (AAC). According to the American Speech-Language and Hearing Association (ASHA), more than two million people in the United States have severe communication disorders that impair their ability to talk. AAC devices that use TTS to create spoken output are used by many of these people to support communication. Historically, AAC users have had access to a relatively small family of generic TTS voices that are neither unique to them nor typically age- or dialect-appropriate. However, advances in TTS technology make it possible to create personalized synthetic voices that capture the unique vocal identity of AAC device users if they are able to record enough speech. This allows patients with neurodegenerative diseases such as ALS to "bank" their voice - that is, to record examples of their speech that can later be used to create a personal TTS voice - before the disease progresses to a point that they can no longer speak. Unfortunately, one major barrier to voice banking, especially for patients who may already be experiencing some difficulty speaking, is the amount of speech needed to create a natural sounding TTS voice that fully captures the vocal identity of the voice banker. To reduce this barrier, this research will combine a type of speech synthesis called parallel formant synthesis that was developed several decades ago, with deep learning computational techniques that allow a computer to learn how to control the parameters of the parallel formant synthesizer to reproduce the speech of a target speaker given examples of the target speaker's speech. A parallel formant synthesizer will be implemented and trained to model speech recorded by voice bankers, and its output will be compared with that of other synthesizers that have been trained with the same speech data. Objective measures of similarity between synthetic and natural utterances, and subjective measures of voice quality and similarity using human listeners, will be used. This will be the first step toward building a parallel formant synthesis-based voice conversion system capable of creating TTS voices from a small number of natural speech samples, and also better able to model the expressive nature of natural speech.Despite advances in TTS technology, there are multiple challenges to the application of this technology for voice banking. Specifically: (a) the amount of speech required (several hours) to create the most natural sounding TTS voices using unit selection or hybrid DNN/unit selection is prohibitive for most voice bankers; (b) existing voice conversion techniques that do not require large amounts of parallel speech from the target talker generally produce speech sounding less natural and less like the target speaker when compared to concatenative synthesis; and (c) both concatenative and statistical parametric techniques produce speech that is only as expressive as the data within the speech corpus from which they have been constructed or trained. Parallel formant synthesis, because it is based explicitly on the perceptually most salient features of natural speech and lends itself to independently modeling laryngeal, suprasegmental, and segmental features should be better able to address all three of these challenges. As proof of concept, a parallel formant synthesis (PFS) vocoder with DNN-based parameter estimation will be implemented. The vocoder will be implemented within the Merlin DNN synthesis framework so that speech output of the PFS system can be directly compared to output generated by the World and MagPhase vocoders. Training will be based on corpora drawn from the same set of 1600 utterances recorded by multiple individuals who have contributed their recordings to the ModelTalker project. The selected target talkers will be balanced for gender and span a wide range of English dialects, but use of speakers with noticeable levels of dysarthria will be avoided. Objective comparisons will be based on Mel-Cepstral Difference (MCD) between synthetic and natural sentence tokens that were not used in training the synthesizers. Subjective measures (Mean Opinion Scores) will be obtained from human listeners via Amazon Mechanical Turk.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Unsupervised Training of a DNN-Based Formant Tracker
基于 DNN 的共振峰跟踪器的无监督训练
DOI: 10.21437/interspeech.2021-1690
发表时间: 2021
期刊: Proceedings of InterSpeech 2021
影响因子: --
作者: [Lilley, Jason, Bunnell, H. Timothy]
通讯作者: Bunnell, H. Timothy
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: