Learning robust speech representation with an articulatory-regularized variational autoencoder

Learning robust speech representation with an articulatory-regularized variational autoencoder
复制标题

使用发音正则化变分自动编码器学习鲁棒的语音表示

DOI:
10.21437/interspeech.2021-1604
复制
发表时间:
2021
期刊:
ArXiv
影响因子:
--
通讯作者:
Thomas Hueber
Thomas Hueber
中科院分区:
--
文献类型:
--
作者:
Marc;Laurent Girin;J. Schwartz;Thomas Hueber

文献摘要

被引文献

相似文献

人们越来越多地认为,人类的言语感知和产生都依赖于发音表征。在本文中,我们研究了这种类型的表示是否可以提高训练用于编码和解码声学语音特征的深度生成模型(这里是变分自动编码器)的性能。首先,我们开发了一个发音模型,能够关联发音参数描述的下巴,舌头,嘴唇和软腭配置声道形状和频谱特征。然后,我们将这些发音参数到一个变分自动编码器应用于频谱特征,通过使用正则化技术,约束部分的潜在空间,以遵循发音轨迹。我们表明,这种发音约束通过减少收敛时间和收敛时的重建损失来改善模型训练,并在语音去噪任务中产生更好的性能。
It is increasingly considered that human speech perception and production both rely on articulatory representations. In this paper, we investigate whether this type of representation could improve the performances of a deep generative model (here a variational autoencoder) trained to encode and decode acoustic speech features. First we develop an articulatory model able to associate articulatory parameters describing the jaw, tongue, lips and velum configurations with vocal tract shapes and spectral features. Then we incorporate these articulatory parameters into a variational autoencoder applied on spectral features by using a regularization technique that constraints part of the latent space to follow articulatory trajectories. We show that this articulatory constraint improves model training by decreasing time to convergence and reconstruction loss at convergence, and yields better performance in a speech denoising task.