Cross-lingual speaker adaptation based on factor analysis using bilingual speech data for HMM-based speech synthesis

Cross-lingual speaker adaptation based on factor analysis using bilingual speech data for HMM-based speech synthesis
复制标题

DOI:
--
复制
发表时间:
2013
期刊:
--
影响因子:
--
通讯作者:
Takenori Yoshimura;Kei Hashimoto;Keiichiro Oura;Yoshihiko Nankaku;K. Tokuda
Takenori Yoshimura;Kei Hashimoto;Keiichiro Oura;Yoshihiko Nankaku;K. Tokuda
中科院分区:
其他
文献类型:
--
作者:
Takenori Yoshimura;Kei Hashimoto;Keiichiro Oura;Yoshihiko Nankaku;K. Tokuda

文献摘要

被引文献

相似文献

针对双语语音数据,提出了一种基于因子分析的跨语言说话人自适应方法。最近针对里昂证券提出了一种基于状态映射的方法。然而,该方法不能仅转换依赖于说话人的特征。此外,目前还没有韵律改编的理论框架。为了解决这些问题,本文提出了一种基于因子分析的双语语音数据CLSA框架。在该方法中,利用双语语音数据,基于单一统计模型,在统一的(最大似然)框架内同时优化表征语言相关声学特征的模型参数和表征说话人特征的因子。这种同时优化有望为期望的说话人特征提供更好的合成语音质量。实验结果表明,该方法比基于状态映射的语音合成方法具有更好的语音合成效果。
This paper proposes a cross-lingual speaker adaptation (CLSA) method based on factor analysis using bilingual speech data. A state-mapping-based method has recently been proposed for CLSA. However, the method cannot transform only speakerdependent characteristics. Furthermore, there is no theoretical framework for adapting prosody. To solve these problems, this paper presents a CLSA framework based on factor analysis using bilingual speech data. In this proposed method, model parameters representing language-dependent acoustic features and factors representing speaker characteristics are simultaneously optimized within a unified (maximum likelihood) framework based on a single statistical model by using bilingual speech data. This simultaneous optimization is expected to deliver a better quality of synthesized speech for the desired speaker characteristics. Experimental results show that the proposed method can synthesize better speech than the state-mappingbased method.