How to improve TTS systems for emotional expressivity

How to improve TTS systems for emotional expressivity
复制标题

如何改进 TTS 系统的情感表达

DOI:
--
复制
发表时间:
2009
期刊:
Interspeech
影响因子:
--
通讯作者:
N. Minematsu
N. Minematsu
中科院分区:
--
文献类型:
--
作者:
Antonio Rui Ferreira Rebordão;S. Masum;K. Hirose;N. Minematsu

文献摘要

被引文献

相似文献

多项实验揭示了当前文本转语音 (TTS) 系统在情感表达方面的弱点。尽管一些 TTS 系统允许基于 XML 的韵律和/或语音变量表示,但很少有出版物考虑在预处理阶段使用智能文本处理来检测情感信息,这些信息可用于定制情感表达所需的参数。本文描述了一种基于情感线索的自动韵律参数化技术。该技术可识别文本中传达的情感信息,并根据其情感内涵,通过 XML 标记分配适当的音高重音和其他韵律参数。此预处理可帮助 TTS 系统生成包含情感线索的合成语音。实验结果令人鼓舞,并表明语音合成中适当的情感表达的可能性。
Several experiments have been carried out that revealed weaknesses of the current Text-To-Speech (TTS) systems in their emotional expressivity. Although some TTS systems allow XML-based representations of prosodic and/or phonetic variables, few publications considered, as a pre-processing stage, the use of intelligent text processing to detect affective information that can be used to tailor the parameters needed for emotional expressivity. This paper describes a technique for an automatic prosodic parameterization based on affective clues. This technique recognizes the affective information conveyed in a text and, accordingly to its emotional connotation, assigns appropriate pitch accents and other prosodic parameters by XML-tagging. This pre-processing assists the TTS system to generate synthesized speech that contains emotional clues. The experimental results are encouraging and suggest the possibility of suitable emotional expressivity in speech synthesis.