An HMM-based speech synthesis system applied to English

An HMM-based speech synthesis system applied to English
复制标题

DOI:
10.1109/wss.2002.1224415
复制
发表时间:
2002
期刊:
Proceedings of 2002 IEEE Workshop on Speech Synthesis, 2002.
影响因子:
--
通讯作者:
Keiichi Tokuda;H. Zen;Alan W. Black
Keiichi Tokuda;H. Zen;Alan W. Black
中科院分区:
其他
文献类型:
--
作者:
Keiichi Tokuda;H. Zen;Alan W. Black

文献摘要

相似文献

本文描述了一个基于隐马尔可夫模型的语音合成系统(HTS),其中语音波形由隐马尔可夫模型本身产生,并将其应用于英语语音合成中,使用Festival的通用语音合成架构。与其他数据驱动的语音合成方法类似,HTS有一个紧凑的语言相关模块:一系列上下文因素。因此,它可以很容易地扩展到其他语言,虽然HTS的第一个版本是为日语实现的。由此产生的HTS运行时引擎的优点是小:小于1兆字节,不包括文本分析部分。此外,HTS可以很容易地改变合成语音的语音特征,通过使用为语音识别开发的说话人自适应技术。基于HMM的方法和其他单元选择方法之间的关系进行了讨论。
This paper describes an HMM-based speech synthesis system (HTS), in which the speech waveform is generated from HMM themselves, and applies it to English speech synthesis using the general speech synthesis architecture of Festival. Similarly to other data-driven speech synthesis approaches, HTS has a compact language dependent module: a list of contextual factors. Thus, it could easily be extended to other languages, though the first version of HTS was implemented for Japanese. The resulting run-time engine of HTS has the advantage of being small: less than 1 Mbyte, excluding text analysis part. Furthermore, HTS can easily change voice characteristics of synthesized speech by using a speaker adaptation technique developed for speech recognition. The relation between the HMM-based approach and other unit selection approaches is also discussed.