Glottal spectral separation for parametric speech synthesis

Glottal spectral separation for parametric speech synthesis
复制标题

DOI:
10.21437/interspeech.2008-176
复制
发表时间:
2008-09
期刊:
Speech Commun.
影响因子:
--
通讯作者:
João P. Cabral;S. Renals;Korin Richmond;J. Yamagishi
João P. Cabral;S. Renals;Korin Richmond;J. Yamagishi
中科院分区:
其他
文献类型:
--
作者:
João P. Cabral;S. Renals;Korin Richmond;J. Yamagishi

文献摘要

被引文献

相似文献

在参数语音合成中使用声门源模型的最大优势是它提供了对语音质量和说话人身份的各个方面进行转换和建模的参数灵活性程度。然而,很少有研究涉及声门源如何影响合成语音的质量。在这里,我们发展了声门频谱分离(GSS)方法,它包括从语音的频谱包络中分离声门源效应。它使我们能够比较低频模型和简单的脉冲激励,使用相同的谱包络来合成语音。知觉评估的结果表明,LF模型的表现明显好于脉冲。通过修改LF参数,GSS方法也被用来成功地将情态语音转换为喘息或紧张的语音。该方法可用于提高基于隐马尔可夫模型的脉冲激励语音合成器的语音质量和信源参数。索引词:声门谱分离,基于隐马尔可夫模型的语音合成,低频模型。
The great advantage of using a glottal source model in parametric speech synthesis is the degree of parametric flexibility it gives to transform and model aspects of voice quality and speaker identity. However, few studies have addressed how the glottal source affects the quality of synthetic speech. Here, we have developed the Glottal Spectral Separation (GSS) method which consists of separating the glottal source effects from the spectral envelope of the speech. It enables us to compare the LF-model with the simple impulse excitation, using the same spectral envelope to synthesize speech. The results of a perceptual evaluation showed that the LF-model clearly outperformed the impulse. The GSS method was also used to successfully transform a modal voice into a breathy or tense voice, by modifying the LF-parameters. The proposed technique could be used to improve the speech quality and source parametrization of HMM-based speech synthesizers, which use an impulse excitation. Index Terms: Glottal Spectral Separation, HMM-based speech synthesis, LF-model.