Duration prediction using multi-level model for GPR-based speech synthesis

Duration prediction using multi-level model for GPR-based speech synthesis
复制标题

DOI:
10.21437/interspeech.2015-369
复制
发表时间:
2015
期刊:
--
影响因子:
--
通讯作者:
Decha Moungsri;Tomoki Koriyama;Takao Kobayashi
Decha Moungsri;Tomoki Koriyama;Takao Kobayashi
中科院分区:
其他
文献类型:
--
作者:
Decha Moungsri;Tomoki Koriyama;Takao Kobayashi

文献摘要

相似文献

本文将基于帧的高斯过程回归 (GPR) 引入泰语语音合成的音素/音节持续时间建模中。 GPR模型旨在利用相应的帧信息来预测帧级声学特征,其中包括每个话语结构单元的相对位置以及诸如声调类型和词性等语言信息。尽管基于 GPR 的预测可以应用于电话持续时间模型,但仅使用电话持续时间模型并不总是足以生成自然的语音。具体来说,在包括泰语在内的某些语言中,音节持续时间会影响对句子结构的感知。在本文中,我们提出了一种使用多级模型的持续时间预测技术,其中包括用于预测的音节和音素级别。在该技术中,首先预测音节持续时间,然后将它们用作音素级模型中的附加上下文来生成用于合成的音素持续时间。客观和主观评估结果表明,基于 GPR 的多级持续时间预测模型建模优于传统的基于 HMM 的语音合成。
This paper introduces frame-based Gaussian process regression (GPR) into phone/syllable duration modeling for Thai speech synthesis. The GPR model is designed for predicting frame-level acoustic features using corresponding frame information, which includes relative position in each unit of utterance structure and linguistic information such as tone type and part of speech. Although the GPR-based prediction can be applied to a phone duration model, the use of phone duration model only is not always sufficient to generate natural sounding speech. Specifically, in some languages including Thai, syllable durations affect the perception of sentence structure. In this paper, we propose a duration prediction technique using a multi-level model which includes syllable and phone levels for prediction. In the technique, first, syllable durations are predicted, and then they are used as additional contexts in phone-level model to generate phone duration for synthesizing. Objective and subjective evaluation results show that GPR-based modeling with multi-level model for duration prediction outperforms the conventional HMM-based speech synthesis.