HMM-based speech synthesiser using the LF-model of the glottal source

HMM-based speech synthesiser using the LF-model of the glottal source
复制标题

DOI:
10.1109/icassp.2011.5947405
复制
发表时间:
2011-05
期刊:
2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
João P. Cabral;S. Renals;J. Yamagishi;Korin Richmond
João P. Cabral;S. Renals;J. Yamagishi;Korin Richmond
中科院分区:
其他
文献类型:
--
作者:
João P. Cabral;S. Renals;J. Yamagishi;Korin Richmond

文献摘要

被引文献

相似文献

在基于HMM的语音合成中引起语音质量恶化的主要因素是使用简单的增量脉冲信号来产生有声语音的激励。本文提出了一种新的方法,在基于HMM的合成器中使用声学声门源模型,而不是传统的脉冲信号。其目标是提高语音质量,更好地建模和转换语音特征。我们已经发现,新的方法减少了bufferk,也提高了韵律建模。感知评估支持这一发现,显示了55.6%的偏好,新系统,与基线。这种改进虽然没有我们最初预期的那么重要,但确实鼓励我们进一步开发所提出的语音合成器。
A major factor which causes a deterioration in speech quality in HMM-based speech synthesis is the use of a simple delta pulse signal to generate the excitation of voiced speech. This paper sets out a new approach to using an acoustic glottal source model in HMM-based synthesisers instead of the traditional pulse signal. The goal is to improve speech quality and to better model and transform voice characteristics. We have found the new method decreases buzziness and also improves prosodic modelling. A perceptual evaluation has supported this finding by showing a 55.6% preference for the new system, as against the baseline. This improvement, while not being as significant as we had initially expected, does encourage us to work on developing the proposed speech synthesiser further.