Acoustic Modeling of Speaking Styles and Emotional Expressions in HMM-Based Speech Synthesis

Acoustic Modeling of Speaking Styles and Emotional Expressions in HMM-Based Speech Synthesis
复制标题

DOI:
10.1093/ietisy/e88-d.3.502
复制
发表时间:
2005-03
影响因子:
0.7
通讯作者:
Junichi Yamagishi;Koji Onishi;T. Masuko;Takao Kobayashi
Junichi Yamagishi;Koji Onishi;T. Masuko;Takao Kobayashi
中科院分区:
计算机科学4区
文献类型:
--
作者:
Junichi Yamagishi;Koji Onishi;T. Masuko;Takao Kobayashi

文献摘要

被引文献

相似文献

本文利用基于hmm的语音合成技术对合成语音中的各种情绪表达和说话风格进行建模。我们展示了两种建模说话风格和情绪表达的方法。在第一种称为风格依赖建模的方法中,每种说话风格和情感表达都是单独建模的。第二种是风格混合建模,将每一种说话风格和情感表达视为一个语境,同时将语音、韵律和语言特征视为一个语境,使用单一的声学模型同时对所有的说话风格和情感表达进行建模。我们选择了四种读语音风格——中性、粗糙、快乐和悲伤——并使用这些风格对上述两种建模方法进行了比较。主观评价测试结果表明,两种建模方法的准确率基本一致,可以合成出与目标语音相似的说话风格和情感表达。在对合成语音风格分类的测试中,使用这两种模型生成的语音样本中有80%以上被判断为与目标风格相似。我们还表明,风格混合建模方法比风格依赖建模方法给出更少的输出和持续时间分布。
This paper describes the modeling of various emotional expressions and speaking styles in synthetic speech using HMM-based speech synthesis. We show two methods for modeling speaking styles and emotional expressions. In the first method called style-dependent modeling, each speaking style and emotional expression is modeled individually. In the second one called style-mixed modeling, each speaking style and emotional expression is treated as one of contexts as well as phonetic, prosodic, and linguistic features, and all speaking styles and emotional expressions are modeled simultaneously by using a single acoustic model. We chose four styles of read speech --- neutral, rough, joyful, and sad --- and compared the above two modeling methods using these styles. The results of subjective evaluation tests show that both modeling methods have almost the same accuracy, and that it is possible to synthesize speech with the speaking style and emotional expression similar to those of the target speech. In a test of classification of styles in synthesized speech, more than 80% of speech samples generated using both the models were judged to be similar to the target styles. We also show that the style-mixed modeling method gives fewer output and duration distributions than the styledependent modeling method.