Cross-lingual speaker adaptation based on factor analysis using bilingual speech data for HMM-based speech synthesis
Cross-lingual speaker adaptation based on factor analysis using bilingual speech data for HMM-based speech synthesis
复制标题
DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
Takenori Yoshimura;Kei Hashimoto;Keiichiro Oura;Yoshihiko Nankaku;K. Tokuda
中科院分区:
文献类型:
--
作者:
Takenori Yoshimura;Kei Hashimoto;Keiichiro Oura;Yoshihiko Nankaku;K. Tokuda
This paper proposes a cross-lingual speaker adaptation (CLSA) method based on factor analysis using bilingual speech data. A state-mapping-based method has recently been proposed for CLSA. However, the method cannot transform only speakerdependent characteristics. Furthermore, there is no theoretical framework for adapting prosody. To solve these problems, this paper presents a CLSA framework based on factor analysis using bilingual speech data. In this proposed method, model parameters representing language-dependent acoustic features and factors representing speaker characteristics are simultaneously optimized within a unified (maximum likelihood) framework based on a single statistical model by using bilingual speech data. This simultaneous optimization is expected to deliver a better quality of synthesized speech for the desired speaker characteristics. Experimental results show that the proposed method can synthesize better speech than the state-mappingbased method.