Modeling of various speaking styles and emotions for HMM-based speech synthesis

Modeling of various speaking styles and emotions for HMM-based speech synthesis
复制标题

DOI:
10.21437/eurospeech.2003-676
复制
发表时间:
2003-09
期刊:
--
影响因子:
--
通讯作者:
J. Yamagishi;Koji Onishi;T. Masuko;Takao Kobayashi
J. Yamagishi;Koji Onishi;T. Masuko;Takao Kobayashi
中科院分区:
其他
文献类型:
--
作者:
J. Yamagishi;Koji Onishi;T. Masuko;Takao Kobayashi

文献摘要

被引文献

相似文献

本文提出了一种基于隐马尔可夫模型的语音合成方法来实现合成语音中各种情感表达和说话风格。我们展示了两种建模说话风格和情绪的方法。在称为“风格依赖建模”的fiRst方法中,每个说话风格和情感都是单独建模的。另一方面,在第二种称为“风格混合建模”的方法中,将说话风格或情感视为一个语境因素,以及语音、韵律和语言因素,所有的说话风格和情感都由一个单一的声学模型同时建模。我们选择了四种风格,即“阅读”、“粗略”、“快乐”和“悲伤”,并比较了使用这四种风格的两种建模方法。主观测试的结果表明,这两种建模方法的性能几乎相同,而且可以合成出与记录语音相似的说话风格和情感的语音。此外,与依赖于风格的建模相比,风格混合建模可以减少输出分布的数目。
This paper presents an approach to realizing various emotional expressions and speaking styles in synthetic speech using HMM-based speech synthesis. We show two methods for modeling speaking styles and emotions. In the first method, called “style dependent modeling,” each speaking style and emotion is individually modeled. On the other hand, in the second method, called “style mixed modeling,” speaking style or emotion is treated as a contextual factor as well as phonetic, prosodic, and linguistic factors, and all speaking styles and emotions are modeled by a single acoustic model simultaneously. We chose four styles, that is, “reading,” “rough,” “joyful,” and “sad,” and compared those two modeling methods using these styles. From the results of subjective tests, it is shown that both modeling meth-ods have almost the same performance, and that it is possible to synthesize speech with similar speaking styles and emotions to those of the recorded speech. In addition, it is also shown that the style mixed modeling can reduce the number of output distributions in comparison with the style dependent modeling.