Speech Synthesis with Various Emotional Expressions and Speaking Styles by Style Interpolation and Morphing

Speech Synthesis with Various Emotional Expressions and Speaking Styles by Style Interpolation and Morphing
复制标题

DOI:
10.1093/ietisy/e88-d.11.2484
复制
发表时间:
2005-11
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
M. Tachibana;J. Yamagishi;T. Masuko;Takao Kobayashi
M. Tachibana;J. Yamagishi;T. Masuko;Takao Kobayashi
中科院分区:
其他
文献类型:
--
作者:
M. Tachibana;J. Yamagishi;T. Masuko;Takao Kobayashi

文献摘要

被引文献

相似文献

本文描述了一种生成具有情感表现力和说话风格可变性的语音的方法。该方法基于基于hmm的语音合成的说话风格和情感表达建模技术。我们首先在基于hmm的语音合成框架中对几种具有代表性的风格进行建模,每种风格都是一种说话风格和/或一种情感表达。然后,为了从具有代表性的语音生成具有中间风格的合成语音,我们使用模型插值技术对具有代表性的风格模型进行插值得到的模型进行合成语音。通过对两种风格组合的插值模型得到的模型,对阅读语音和合成语音中具有代表性的中性、愉悦、悲伤和粗糙四种风格进行主观评价测试,对风格插值技术进行了评价。结果表明,由插值模型合成的语音具有介于两种代表性语音之间的风格。此外,我们可以通过改变中性和其他代表性风格之间的插值比例来控制合成语音中说话风格或情绪的表达程度。我们还表明,我们可以在语音合成中实现风格变形,即通过逐渐改变插值比,从一种代表风格平滑地转换到另一种代表风格。
This paper describes an approach to generating speech with emotional expressivity and speaking style variability. The approach is based on a speaking style and emotional expression modeling technique for HMM-based speech synthesis. We first model several representative styles, each of which is a speaking style and/or an emotional expression, in an HMM-based speech synthesis framework. Then, to generate synthetic speech with an intermediate style from representative ones, we synthesize speech from a model obtained by interpolating representative style models using a model interpolation technique. We assess the style interpolation technique with subjective evaluation tests using four representative styles, i.e., neutral, joyful, sad, and rough in read speech and synthesized speech from models obtained by interpolating models for all combinations of two styles. The results show that speech synthesized from the interpolated model has a style in between the two representative ones. Moreover, we can control the degree of expressivity for speaking styles or emotions in synthesized speech by changing the interpolation ratio in interpolation between neutral and other representative styles. We also show that we can achieve style morphing in speech synthesis, namely, changing style smoothly from one representative style to another by gradually changing the interpolation ratio.