Deep Speech Synthesis from Articulatory Representations

Deep Speech Synthesis from Articulatory Representations
复制标题

DOI:
10.21437/interspeech.2022-10892
复制
发表时间:
2022-09
期刊:
--
影响因子:
--
通讯作者:
Peter Wu;Shinji Watanabe;L. Goldstein;A. Black;G. Anumanchipalli
Peter Wu;Shinji Watanabe;L. Goldstein;A. Black;G. Anumanchipalli
中科院分区:
其他
文献类型:
--
作者:
Peter Wu;Shinji Watanabe;L. Goldstein;A. Black;G. Anumanchipalli

文献摘要

被引文献

相似文献

在发音合成任务中,语音是从包含关于人类声道物理行为的信息的输入特征合成的。这项任务为语音合成研究提供了一个很有前途的方向,因为发音空间是紧凑的、平滑的和可解释的。目前的工作强调了深度学习模型执行发音合成的潜力。然而,目前尚不清楚这些模型能否达到人类语音生成系统的效率和保真度。为了帮助弥补这一差距,我们提出了一种时域发音合成方法,并通过电磁关节成像(EMA)和合成发音特征输入验证了其有效性。我们的模型计算效率很高,在EMA-to-Speech任务中获得了18.5%的转录单词错误率(WER),与以前的工作相比提高了11.6%。通过内插实验,突出了该方法的泛化能力和可解释性。
In the articulatory synthesis task, speech is synthesized from input features containing information about the physical behavior of the human vocal tract. This task provides a promising direction for speech synthesis research, as the articulatory space is compact, smooth, and interpretable. Current works have highlighted the potential for deep learning models to perform articulatory synthesis. However, it remains unclear whether these models can achieve the efficiency and fidelity of the human speech production system. To help bridge this gap, we propose a time-domain articulatory synthesis methodology and demonstrate its efficacy with both electromagnetic articulography (EMA) and synthetic articulatory feature inputs. Our model is computationally efficient and achieves a transcription word error rate (WER) of 18.5% for the EMA-to-speech task, yielding an improvement of 11.6% compared to prior work. Through interpolation experiments, we also highlight the generalizability and interpretability of our approach.