From HMMS to DNNS: Where do the improvements come from?

From HMMS to DNNS: Where do the improvements come from?
复制标题

DOI:
10.1109/icassp.2016.7472730
复制
发表时间:
2016-03
期刊:
2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
O. Watts;G. Henter;Thomas Merritt;Zhizheng Wu;Simon King
O. Watts;G. Henter;Thomas Merritt;Zhizheng Wu;Simon King
中科院分区:
其他
文献类型:
--
作者:
O. Watts;G. Henter;Thomas Merritt;Zhizheng Wu;Simon King

文献摘要

被引文献

相似文献

深度神经网络(DNN)作为统计参数合成系统中决策树和隐马尔可夫模型(HMM)的替代方法,近年来成为文语转换研究的热点。已经有性能改进的报告;然而,被评估的系统的配置使得无法判断这种改进在多大程度上是由于新的机器学习方法,以及在多大程度上是由于系统的其他新方面。具体地说,尽管基于HMM的系统中的决策树通常在状态级别上操作,并且使用单独的树来处理单独的声学流,但是大多数基于DNN的系统被训练为在声学帧的级别上同时对所有流进行预测。本文通过建立一个一次只有一个因素变化的系统连续体来隔离三个因素(机器学习方法;状态预测与帧预测;单独流预测与组合流预测)的影响。我们发现,用DNN代替决策树和从状态级预测转移到帧级预测都显著提高了听者对系统产生的合成语音的自然度评级。从分流预测转换为合流预测没有发现任何改善。
Deep neural networks (DNNs) have recently been the focus of much text-to-speech research as a replacement for decision trees and hidden Markov models (HMMs) in statistical parametric synthesis systems. Performance improvements have been reported; however, the configuration of systems evaluated makes it impossible to judge how much of the improvement is due to the new machine learning methods, and how much is due to other novel aspects of the systems. Specifically, whereas the decision trees in HMM-based systems typically operate at the state-level, and separate trees are used to handle separate acoustic streams, most DNN-based systems are trained to make predictions simultaneously for all streams at the level of the acoustic frame. This paper isolates the influence of three factors (machine learning method; state vs. frame predictions; separate vs. combined stream predictions) by building a continuum of systems along which only a single factor is varied at a time. We find that replacing decision trees with DNNs and moving from state-level to frame-level predictions both significantly improve listeners' naturalness ratings of synthetic speech produced by the systems. No improvement is found to result from switching from separate-stream to combined-stream predictions.