Speaker and style adaptation using average voice model for style control in HMM-based speech synthesis

Speaker and style adaptation using average voice model for style control in HMM-based speech synthesis
复制标题

DOI:
10.1109/icassp.2008.4518689
复制
发表时间:
2008-05
期刊:
2008 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
通讯作者:
M. Tachibana;Shinsuke Izawa;Takashi Nose;Takao Kobayashi
M. Tachibana;Shinsuke Izawa;Takashi Nose;Takao Kobayashi
中科院分区:
其他
文献类型:
--
作者:
M. Tachibana;Shinsuke Izawa;Takashi Nose;Takao Kobayashi

文献摘要

被引文献

相似文献

我们提出了一种合成任意目标说话人的语音所需风格表达的技术。在基于mllr的多重回归隐马尔可夫模型(MRHSMM)的说话人自适应技术中,合成语音的质量很大程度上取决于从特定的源说话人数据中训练的初始MRHSMM,并且并不总是能够用给定的目标说话人的声音合成自然发音的语音。为了克服这个问题,我们从一个平均语音模型中同时进行说话人和风格的适应。实验结果表明,该方法比仅使用说话人自适应的传统方法提供了更自然的语音。
We propose a technique for synthesizing speech with desired style expressivity of an arbitrary target speaker's voice. In an MLLR-based speaker adaptation technique for multiple regression hidden semi-Markov model (MRHSMM), the quality of synthesized speech crucially depends on the initial MRHSMM trained from a certain source speaker's data and it is not always possible to synthesize natural sounding speech with a given target speaker's voice. To overcome this problem, we perform simultaneous adaptation of speaker and style from an average voice model. Experimental results show that the proposed technique provides more natural sounding speech than the conventional one with speaker adaptation only.