Prosody generation using frame-based Gaussian process regression and classification for statistical parametric speech synthesis

Prosody generation using frame-based Gaussian process regression and classification for statistical parametric speech synthesis
复制标题

DOI:
10.1109/icassp.2015.7178908
复制
发表时间:
2015-04
期刊:
2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Tomoki Koriyama;Takao Kobayashi
Tomoki Koriyama;Takao Kobayashi
中科院分区:
其他
文献类型:
--
作者:
Tomoki Koriyama;Takao Kobayashi

文献摘要

相似文献

本文提出了一种新的基于高斯过程回归和分类的F0轮廓和语音时长模型(GPR和GPC),用于统计参数语音合成。虽然基于帧的探地雷达的使用在以往的研究中显示了谱特征建模的有效性,但由于核函数仅用于语音信息,因此将探地雷达应用于韵律特征,即F0和音节时长,还没有得到充分的研究。因此,在本文中,我们提出了一种适用于音节、音节和重音短语等多个单元的核函数。该核函数基于重音短语开头等时间声学事件,并利用目标帧与事件的相对位置作为核函数。客观测试和主观测试的实验结果表明,与基于HMM的语音合成相比,基于GPR/GPC的F0和时长建模提高了声学特征的预测精度。
This paper proposes novel models of F0 contours and phone durations using Gaussian process regression and classification (GPR and GPC) for statistical parametric speech synthesis. Although the use of frame-based GPR has shown the effectiveness of spectral feature modeling in previous studies, the application of GPR to prosodic features, i.e., F0 and phone duration, was not investigated sufficiently because the kernel function was designed for phonetic information only. In this paper, therefore, we propose a kernel function available for multiple units such as syllables, moras, and accent phrases. The proposed kernel function is based on temporal acoustic events like the beginning of accent phrase and the relative position between the target frame and the event is utilized for the kernel function. Experimental results of objective and subjective tests show that the GPR/GPC-based F0 and duration modeling improves the prediction accuracy of acoustic features compared with HMM-based speech synthesis.