Deep Speech Synthesis from MRI-Based Articulatory Representations

Deep Speech Synthesis from MRI-Based Articulatory Representations
复制标题

DOI:
10.21437/interspeech.2023-2316
复制
发表时间:
2023-07
期刊:
--
影响因子:
--
通讯作者:
Peter Wu;Tingle Li;Yijingxiu Lu;Yubin Zhang;Jiachen Lian;A. Black;L. Goldstein;Shinji Watanabe;G. Anumanchipalli
Peter Wu;Tingle Li;Yijingxiu Lu;Yubin Zhang;Jiachen Lian;A. Black;L. Goldstein;Shinji Watanabe;G. Anumanchipalli
中科院分区:
其他
文献类型:
--
作者:
Peter Wu;Tingle Li;Yijingxiu Lu;Yubin Zhang;Jiachen Lian;A. Black;L. Goldstein;Shinji Watanabe;G. Anumanchipalli

文献摘要

被引文献

相似文献

在本文中,我们研究发音合成,语音合成方法,使用人类声道信息,提供了一种方法来开发高效,通用和可解释的合成器。虽然最近的进展已经使可理解的发音合成使用电磁关节描记术(EMA),这些方法缺乏关键的发音信息,如兴奋和鼻音,限制了泛化能力。为了弥合这一差距,我们提出了一个替代的基于MRI的功能集,涵盖了更广泛的发音空间比EMA。我们还引入了归一化和去噪程序,以增强在MRI数据上训练的深度学习方法的通用性。此外,我们提出了一个MRI语音模型,提高了计算效率和语音保真度。最后,通过一系列的消融,我们表明,建议的MRI表示比EMA更全面,并确定最合适的MRI特征子集的发音合成。
In this paper, we study articulatory synthesis, a speech synthesis method using human vocal tract information that offers a way to develop efficient, generalizable and interpretable synthesizers. While recent advances have enabled intelligible articulatory synthesis using electromagnetic articulography (EMA), these methods lack critical articulatory information like excitation and nasality, limiting generalization capabilities. To bridge this gap, we propose an alternative MRI-based feature set that covers a much more extensive articulatory space than EMA. We also introduce normalization and denoising procedures to enhance the generalizability of deep learning methods trained on MRI data. Moreover, we propose an MRI-to-speech model that improves both computational efficiency and speech fidelity. Finally, through a series of ablations, we show that the proposed MRI representation is more comprehensive than EMA and identify the most suitable MRI feature subset for articulatory synthesis.