An autoregressive recurrent mixture density network for parametric speech synthesis

An autoregressive recurrent mixture density network for parametric speech synthesis
复制标题

DOI:
10.1109/icassp.2017.7953087
复制
发表时间:
2017-06
期刊:
2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Xin Wang;Shinji Takaki;J. Yamagishi
Xin Wang;Shinji Takaki;J. Yamagishi
中科院分区:
其他
文献类型:
--
作者:
Xin Wang;Shinji Takaki;J. Yamagishi

文献摘要

相似文献

基于神经网络的生成模型,例如混合密度网络,是语音合成的潜在解决方案。在本文中,我们遵循这条路径,提出了一种包含可训练自回归模型的循环混合密度网络。合并自回归模型的优点是可以在不使用传统动态特征的情况下对声学特征轨迹内的时间依赖性进行建模。更有趣的是,实验表明,这种自回归模型在训练阶段学习成为一个强调目标声学特征轨迹高频分量的滤波器。在合成阶段,它增强了生成的特征轨迹的低频分量,从而增加了它们的全局方差。实验结果表明,当任何模型中未使用动态特征时,所提出的模型在训练数据上获得了更高的似然性,并且生成的语音质量比其他模型更好。
Neural-network-based generative models, such as mixture density networks, are potential solutions for speech synthesis. In this paper we follow this path and propose a recurrent mixture density network that incorporates a trainable autoregressive model. An advantage of incorporating an autoregressive model is that the time dependency within acoustic feature trajectories can be modeled without using the conventional dynamic features. More interestingly, experiments show that this autoregressive model learns to be a filter that emphasizes the high frequency components of the target acoustic feature trajectories in the training stage. In the synthesis stage, it boosts the low frequency components of the generated feature trajectories and hence increases their global variance. Experimental results show that the proposed model achieved higher likelihood on the training data and generated speech with better quality than other models when dynamic features were not utilized in any model.