Expressive Speech Synthesis: Past, Present, and Possible Futures

Expressive Speech Synthesis: Past, Present, and Possible Futures
复制标题

DOI:
10.1007/978-1-84800-306-4_7
复制
发表时间:
2009
期刊:
--
影响因子:
--
通讯作者:
M. Schröder
M. Schröder
中科院分区:
其他
文献类型:
--
作者:
M. Schröder

文献摘要

被引文献

相似文献

过去 20 年来,为合成语音添加表现力的方法发生了很大变化。早期的系统,包括共振峰和双音素系统,一直专注于“显式控制”模型。早期的单位选择系统采用了“回放”的方式。目前,人们正在寻求各种方法来提高表达的灵活性,同时保持最先进系统的质量,其中包括统计参数语音合成中的新“隐式控制”范式,它通过在不同表达数据库上训练的统计模型之间进行组合和插值来控制表达能力。本章概述了过去和现在的方法,并探讨了未来可能的发展。
Approaches towards adding expressivity to synthetic speech have changed considerably over the last 20 years. Early systems, including formant and diphone systems, have been focused around “explicit control” models; early unit selection systems have adopted a “playback” approach. Currently, various approaches are being pursued to increase the flexibility in expression while maintaining the quality of state-of-the-art systems, among them a new “implicit control” paradigm in statistical parametric speech synthesis, which provides control over expressivity by combining and interpolating between statistical models trained on different expressive databases. The present chapter provides an overview of the past and present approaches, and ventures a look into possible future developments.