Speaker and style adaptation using average voice model for style control in HMM-based speech synthesis
Speaker and style adaptation using average voice model for style control in HMM-based speech synthesis
复制标题
DOI:
10.1109/icassp.2008.4518689
复制
发表时间:
2008-05
期刊:
影响因子:
--
通讯作者:
M. Tachibana;Shinsuke Izawa;Takashi Nose;Takao Kobayashi
中科院分区:
文献类型:
--
作者:
M. Tachibana;Shinsuke Izawa;Takashi Nose;Takao Kobayashi
We propose a technique for synthesizing speech with desired style expressivity of an arbitrary target speaker's voice. In an MLLR-based speaker adaptation technique for multiple regression hidden semi-Markov model (MRHSMM), the quality of synthesized speech crucially depends on the initial MRHSMM trained from a certain source speaker's data and it is not always possible to synthesize natural sounding speech with a given target speaker's voice. To overcome this problem, we perform simultaneous adaptation of speaker and style from an average voice model. Experimental results show that the proposed technique provides more natural sounding speech than the conventional one with speaker adaptation only.