A perceptual investigation of wavelet-based decomposition of f0 for text-to-speech synthesis
A perceptual investigation of wavelet-based decomposition of f0 for text-to-speech synthesis
复制标题
DOI:
10.21437/interspeech.2015-368
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
M. Ribeiro;J. Yamagishi;R. Clark
中科院分区:
文献类型:
--
作者:
M. Ribeiro;J. Yamagishi;R. Clark
The Continuous Wavelet Transform (CWT) has been re- cently proposed to model f0 in the context of speech synthe- sis. It was shown that systems using signal decomposition with the CWT tend to outperform systems that model the signal di- rectly. The f0 signal is typically decomposed into various scales of differing frequency. In these experiments, we reconstruct f0 with selected frequencies and ask native listeners to judge the naturalness of synthesized utterances with respect to natural speech. Results indicate that HMM-generated f0 is compara- ble to the CWT low frequencies, suggesting it mostly generates utterances with neutral intonation. Middle frequencies achieve very high levels of naturalness, while very high frequencies are mostly noise.