Parametric speech synthesis based on Gaussian process regression using global variance and hyperparameter optimization

Parametric speech synthesis based on Gaussian process regression using global variance and hyperparameter optimization
复制标题

基于使用全局方差和超参数优化的高斯过程回归的参数语音合成

DOI:
10.1109/icassp.2014.6854319
复制
发表时间:
2014
期刊:
Proceedings of 2014 IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
Takao Kobayashi
Takao Kobayashi
中科院分区:
--
文献类型:
--
作者:
Tomoki Koriyama;Takashi Nose;Takao Kobayashi

文献摘要

相似文献

本文研究了基于高斯过程回归的统计语音合成方法的两个问题。尽管基于gp的语音合成在生成频谱参数方面比基于hmm的语音合成具有更高的性能,但仍存在一些问题。本文将全局方差(global variance, GV)特征引入到参数生成中,克服了过度平滑问题。此外,为了根据实际数据使用合适的核函数,我们提出了一种基于em的核超参数优化技术。客观和主观评价结果表明,利用梯度向量和超参数估计增强了光谱特征生成的性能。
This paper examines two issues of a statistical speech synthesis approach based Gaussian process (GP) regression. Although GP-based speech synthesis can give higher performance in generating spectral parameters than the HMM-based one, a number of issues still remain. In this paper, we incorporate global variance (GV) feature to overcome over-smoothing problem into the parameter generation. Furthermore, in order to utilize an appropriate kernel function in accordance with actual data, we propose an EM-based kernel hyperparameter optimization technique. Objective and subjective evaluation results show that using GV and hyperparameter estimation enhanced the performance in spectral feature generation.