Statistical nonparametric speech synthesis using sparse Gaussian processes

Statistical nonparametric speech synthesis using sparse Gaussian processes
复制标题

DOI:
10.21437/interspeech.2013-121
复制
发表时间:
2013
期刊:
--
影响因子:
--
通讯作者:
Tomoki Koriyama;Takashi Nose;Takao Kobayashi
Tomoki Koriyama;Takashi Nose;Takao Kobayashi
中科院分区:
其他
文献类型:
--
作者:
Tomoki Koriyama;Takashi Nose;Takao Kobayashi

文献摘要

相似文献

提出了一种基于稀疏高斯过程回归(GPR)的统计非参数语音合成技术。在我们之前的研究中,我们提出了基于GPR的语音合成,其中合成单元的每一帧由高斯过程的回归建模。初步的实验合成几个音素,包括元音和conso-contrast显示了该技术的潜力。在本文中,以前的工作扩展到使用稀疏GP和上下文修改的整句语音合成。具体地说,基于集群的稀疏高斯过程,如局部GP和部分独立条件(PIC)近似作为一种计算上可行的方法进行了研究。此外,帧级上下文被扩展为不仅包括来自当前音素的位置上下文,还包括相邻音素,以生成平滑变化的语音参数。客观和主观评价结果表明,该技术优于基于HMM的语音合成与最小生成误差训练。
This paper proposes a statistical nonparametric speech synthesis technique based on a sparse Gaussian process regression (GPR). In our previous study, we proposed GPR-based speech synthesis where each frame of synthesis units is modeled by a regression of Gaussian processes. Preliminary experiments of synthesizing several phones including both vowels and conso-nants showed a potential of the technique. In this paper, the previous work is extended to full-sentence speech synthesis using sparse GPs and context modification. Specifically, cluster-based sparse Gaussian processes such as local GPs and partially independent conditional (PIC) approximation are examined as a computationally feasible approach. Moreover, frame-level context is extended to include not only a position context from a current phone but also adjacent phones to generate smoothly changing speech parameters. Objective and subjective evaluation results show that the proposed technique outperforms the HMM-based speech synthesis with minimum generation error training.