Evaluation of Expressive Speech Synthesis With Voice Conversion and Copy Resynthesis Techniques

Evaluation of Expressive Speech Synthesis With Voice Conversion and Copy Resynthesis Techniques
复制标题

DOI:
10.1109/tasl.2010.2041113
复制
发表时间:
2010-07
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
O. Türk;M. Schröder
O. Türk;M. Schröder
中科院分区:
其他
文献类型:
--
作者:
O. Türk;M. Schröder

文献摘要

被引文献

相似文献

生成富有表现力的合成语音需要精心设计的数据库,其中包含足够数量的富有表现力的语音材料。本文研究语音转换和修改技术,以减少数据库的收集和处理工作,同时保持可接受的质量和自然。在析因设计中,我们研究了语音质量和韵律的相对贡献,以及由各自的信号处理步骤引入的失真量。我们的开源和模块化文本到语音(TTS)框架玛丽中的单元选择引擎扩展了使用基于GMM的预测或声道复制再合成的语音质量转换。然后将这些算法与各种韵律复制再合成方法交叉组合。整个表达性语音生成过程用作TTS输出的后处理步骤,以将中性合成语音转换为攻击性,愉快或沮丧的语音。语音质量和韵律转换算法的交叉组合进行了比较,在听力测试的感知表达风格和质量。结果表明,在识别和自然性之间存在一个折衷。语音质量和韵律的组合建模导致以最低的自然度评级为代价的最佳识别分数。语音质量和韵律的细节,保留了复制合成,有助于更好地识别相比,近似模型。
Generating expressive synthetic voices requires carefully designed databases that contain sufficient amount of expressive speech material. This paper investigates voice conversion and modification techniques to reduce database collection and processing efforts while maintaining acceptable quality and naturalness. In a factorial design, we study the relative contributions of voice quality and prosody as well as the amount of distortions introduced by the respective signal manipulation steps. The unit selection engine in our open source and modular text-to-speech (TTS) framework MARY is extended with voice quality transformation using either GMM-based prediction or vocal tract copy resynthesis. These algorithms are then cross-combined with various prosody copy resynthesis methods. The overall expressive speech generation process functions as a postprocessing step on TTS outputs to transform neutral synthetic speech into aggressive, cheerful, or depressed speech. Cross-combinations of voice quality and prosody transformation algorithms are compared in listening tests for perceived expressive style and quality. The results show that there is a tradeoff between identification and naturalness. Combined modeling of both voice quality and prosody leads to the best identification scores at the expense of lowest naturalness ratings. The fine detail of both voice quality and prosody, as preserved by the copy synthesis, did contribute to a better identification as compared to the approximate models.