Parametric speech synthesis using local and global sparse Gaussian processes

Parametric speech synthesis using local and global sparse Gaussian processes
复制标题

DOI:
10.1109/mlsp.2014.6958921
复制
发表时间:
2014-11
期刊:
2014 IEEE International Workshop on Machine Learning for Signal Processing (MLSP)
影响因子:
--
通讯作者:
Tomoki Koriyama;Takashi Nose;Takao Kobayashi
Tomoki Koriyama;Takashi Nose;Takao Kobayashi
中科院分区:
其他
文献类型:
--
作者:
Tomoki Koriyama;Takashi Nose;Takao Kobayashi

文献摘要

相似文献

本文描述了高斯过程回归(GPR)在参数语音合成中的应用。GPR使我们能够预测合成语音参数,直接利用训练语音数据的样本,而无需将训练数据的声学特征转换为太少的模型参数,这要归功于非参数贝叶斯回归。然而,GPR固有地需要高计算成本和资源。在本文中,为了缓解这个问题,我们将局部和全局稀疏高斯过程近似的统计语音合成框架,并通过实验研究计算成本和语音合成性能之间的权衡。此外,我们研究了如何选择用于稀疏GP近似的伪数据集。
This paper describes an application of Gaussian process regression (GPR) to parametric speech synthesis. GPR enables us to predict synthetic speech parameters by utilizing exemplars of training speech data directly without converting the acoustic features of training data into too small number of model parameters thanks to nonparametric Bayesian regression. However, GPR inherently requires high computational cost and resources. In this paper, to alleviate this problem, we incorporate local and global sparse Gaussian process approximation into the statistical speech synthesis framework, and investigate trade-off between computational cost and speech synthesis performance through experiments. Moreover, we examine the way of choosing pseudo data set used for the sparse GP approximation.