A perceptual investigation of wavelet-based decomposition of f0 for text-to-speech synthesis

A perceptual investigation of wavelet-based decomposition of f0 for text-to-speech synthesis
复制标题

DOI:
10.21437/interspeech.2015-368
复制
发表时间:
2015
期刊:
--
影响因子:
--
通讯作者:
M. Ribeiro;J. Yamagishi;R. Clark
M. Ribeiro;J. Yamagishi;R. Clark
中科院分区:
其他
文献类型:
--
作者:
M. Ribeiro;J. Yamagishi;R. Clark

文献摘要

被引文献

相似文献

连续小波变换(CWT)是近年来提出的语音合成中的f0模型。结果表明,使用信号分解与连续小波变换的系统往往优于直接建模信号的系统。f0信号通常被分解为不同频率的各种尺度。在这些实验中,我们重建f0与选定的频率,并要求本地听众判断自然语音合成的话语相对于自然。结果表明,HMM生成的f0与CWT低频相当,这表明它主要生成中性语调的话语。中频达到非常高的自然水平,而非常高的频率主要是噪音。
The Continuous Wavelet Transform (CWT) has been re- cently proposed to model f0 in the context of speech synthe- sis. It was shown that systems using signal decomposition with the CWT tend to outperform systems that model the signal di- rectly. The f0 signal is typically decomposed into various scales of differing frequency. In these experiments, we reconstruct f0 with selected frequencies and ask native listeners to judge the naturalness of synthesized utterances with respect to natural speech. Results indicate that HMM-generated f0 is compara- ble to the CWT low frequencies, suggesting it mostly generates utterances with neutral intonation. Middle frequencies achieve very high levels of naturalness, while very high frequencies are mostly noise.