Statistical Parametric Speech Synthesis Based on Gaussian Process Regression

Statistical Parametric Speech Synthesis Based on Gaussian Process Regression
复制标题

DOI:
10.1109/jstsp.2013.2283461
复制
发表时间:
2014-04
影响因子:
7.5
通讯作者:
Tomoki Koriyama;Takashi Nose;Takao Kobayashi
Tomoki Koriyama;Takashi Nose;Takao Kobayashi
中科院分区:
工程技术1区
文献类型:
--
作者:
Tomoki Koriyama;Takashi Nose;Takao Kobayashi

文献摘要

相似文献

提出了一种基于高斯过程回归(GPR)的统计参数语音合成技术。GPR模型的目的是直接预测帧级的声学特征,从相应的信息框架上下文,从语言信息中获得。框架上下文包括当前框架在电话中的相对位置和发音信息,并用作GPR中的解释变量。在这里,我们引入基于聚类的稀疏高斯过程(GP),即,局部GP和部分独立条件(PIC)近似,以减少计算成本。孤立的音素合成和整句连续语音合成的实验结果表明,所提出的基于GPR的技术,没有动态功能略优于传统的隐马尔可夫模型(HMM)为基础的语音合成,使用最小生成误差训练与动态功能。
This paper proposes a statistical parametric speech synthesis technique based on Gaussian process regression (GPR). The GPR model is designed for directly predicting frame-level acoustic features from corresponding information on frame context that is obtained from linguistic information. The frame context includes the relative position of the current frame within the phone and articulatory information and is used as the explanatory variable in GPR. Here, we introduce cluster-based sparse Gaussian processes (GPs), i.e., local GPs and partially independent conditional (PIC) approximation, to reduce the computational cost. The experimental results for both isolated phone synthesis and full-sentence continuous speech synthesis revealed that the proposed GPR-based technique without dynamic features slightly outperformed the conventional hidden Markov model (HMM)-based speech synthesis using minimum generation error training with dynamic features.