Attributing modelling errors in HMM synthesis by stepping gradually from natural to modelled speech

Attributing modelling errors in HMM synthesis by stepping gradually from natural to modelled speech
复制标题

通过逐步从自然语音到建模语音来归因 HMM 合成中的建模错误

DOI:
10.1109/icassp.2015.7178766
复制
发表时间:
2015
期刊:
2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Simon King
Simon King
中科院分区:
--
文献类型:
--
作者:
Thomas Merritt;Javier Latorre;Simon King

文献摘要

参考文献

被引文献

相似文献

即使是最好的统计参数语音合成系统也不能达到好的单元选择的自然性。我们调查了造成这种情况的可能原因。通过构造介于自然语音和完整HMM合成系统输出之间的语音信号,我们研究了建模的各种效果。我们操纵了频谱参数的时间平稳性和方差来创建刺激,然后将这些刺激与自然语音和声编码语音一起提供给听众,以及来自完全基于HMM的文本到语音系统的输出和来自理想化的伪HMM的输出。除自然波形外,所有语音信号都是使用采用两种流行的谱参数之一的声码器创建的:Mel-Cepstra或Mel-Line频谱对。听者做出“相同或不同”的成对判断,我们从这些判断中生成了一个使用多维尺度的感知地图。我们得出了关于HMM合成的哪些方面限制了合成语音的自然度的结论。
Even the best statistical parametric speech synthesis systems do not achieve the naturalness of good unit selection. We investigated possible causes of this. By constructing speech signals that lie in between natural speech and the output from a complete HMM synthesis system, we investigated various effects of modelling. We manipulated the temporal smoothness and the variance of the spectral parameters to create stimuli, then presented these to listeners alongside natural and vocoded speech, as well as output from a full HMM-based text-to-speech system and from an idealised `pseudo-HMM'. All speech signals, except the natural waveform, were created using vocoders employing one of two popular spectral parameterisations: Mel-Cepstra or Mel-Line Spectral Pairs. Listeners made `same or different' pairwise judgements, from which we generated a perceptual map using Multidimensional Scaling. We draw conclusions about which aspects of HMM synthesis are limiting the naturalness of the synthetic speech.
DOI: 10.21437/blizzard.2008-1
发表时间: 2008-09
期刊: The Blizzard Challenge 2008
影响因子: --
作者:
Simon King;R. Clark;C. Mayo;Vasilis Karaiskos
通讯作者: Simon King;R. Clark;C. Mayo;Vasilis Karaiskos