Statistical parametric speech synthesis using deep neural networks

Statistical parametric speech synthesis using deep neural networks
复制标题

DOI:
10.1109/icassp.2013.6639215
复制
发表时间:
2013-05
期刊:
2013 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
通讯作者:
H. Zen;A. Senior;M. Schuster
H. Zen;A. Senior;M. Schuster
中科院分区:
其他
文献类型:
--
作者:
H. Zen;A. Senior;M. Schuster

文献摘要

被引文献

相似文献

统计参数语音合成的常规方法通常使用决策树聚类上下文相关隐马尔可夫模型(HMRM)来表示给定文本的语音参数的概率密度。从概率密度生成语音参数以最大化它们的输出概率,然后从所生成的参数重构语音波形。这种方法相当有效,但也有一些局限性,例如,决策树对复杂的上下文依赖关系建模效率低下。本文研究了一种基于深度神经网络(DNN)的替代方案。输入文本与其声学实现之间的关系由DNN建模。DNN的使用可以解决传统方法的一些限制。实验结果表明,基于DNN的系统优于基于HMM的系统具有相似的参数数量。
Conventional approaches to statistical parametric speech synthesis typically use decision tree-clustered context-dependent hidden Markov models (HMMs) to represent probability densities of speech parameters given texts. Speech parameters are generated from the probability densities to maximize their output probabilities, then a speech waveform is reconstructed from the generated parameters. This approach is reasonably effective but has a couple of limitations, e.g. decision trees are inefficient to model complex context dependencies. This paper examines an alternative scheme that is based on a deep neural network (DNN). The relationship between input texts and their acoustic realizations is modeled by a DNN. The use of the DNN can address some limitations of the conventional approach. Experimental results show that the DNN-based systems outperformed the HMM-based systems with similar numbers of parameters.