Robust Speaker-Adaptive HMM-Based Text-to-Speech Synthesis

Robust Speaker-Adaptive HMM-Based Text-to-Speech Synthesis
复制标题

DOI:
10.1109/tasl.2009.2016394
复制
发表时间:
2009-08
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
J. Yamagishi;Takashi Nose;H. Zen;Zhenhua Ling;T. Toda;K. Tokuda;Simon King;S. Renals
J. Yamagishi;Takashi Nose;H. Zen;Zhenhua Ling;T. Toda;K. Tokuda;Simon King;S. Renals
中科院分区:
其他
文献类型:
--
作者:
J. Yamagishi;Takashi Nose;H. Zen;Zhenhua Ling;T. Toda;K. Tokuda;Simon King;S. Renals

文献摘要

被引文献

相似文献

本文描述了一种基于说话人自适应HMM的语音合成系统。这个名为ldquots-2007的新系统采用了说话人自适应(CSMAPLR+MAP)、特征空间自适应训练、混合性别建模和使用CSMAPLR变换的全协方差建模,以及在我们以前的系统中被证明有效的其他几种技术。主观评价结果表明,在实际语音数据量下,新系统生成的合成语音质量明显优于依赖于说话人的方法,并且即使在有大量语音数据的情况下,它也与依赖于说话人的方法相比较。此外,与几种语音合成技术的比较研究表明,新系统非常健壮:它能够从不太理想的语音数据中构建语音,并合成高质量的语音,即使是对于域外句子。
This paper describes a speaker-adaptive HMM-based speech synthesis system. The new system, called ldquoHTS-2007,rdquo employs speaker adaptation (CSMAPLR+MAP), feature-space adaptive training, mixed-gender modeling, and full-covariance modeling using CSMAPLR transforms, in addition to several other techniques that have proved effective in our previous systems. Subjective evaluation results show that the new system generates significantly better quality synthetic speech than speaker-dependent approaches with realistic amounts of speech data, and that it bears comparison with speaker-dependent approaches even when large amounts of speech data are available. In addition, a comparison study with several speech synthesis techniques shows the new system is very robust: It is able to build voices from less-than-ideal speech data and synthesize good-quality speech even for out-of-domain sentences.