A style control technique for speech synthesis using multiple regression HSMM

A style control technique for speech synthesis using multiple regression HSMM
复制标题

DOI:
10.21437/interspeech.2006-388
复制
发表时间:
2006
期刊:
--
影响因子:
--
通讯作者:
Takashi Nose;J. Yamagishi;Takao Kobayashi
Takashi Nose;J. Yamagishi;Takao Kobayashi
中科院分区:
其他
文献类型:
--
作者:
Takashi Nose;J. Yamagishi;Takao Kobayashi

文献摘要

被引文献

相似文献

本文提出了一种直观地控制合成语音的说话风格和情感表达的程度或强度的技术。传统的基于多元回归隐马尔可夫模型(MRHMM)的风格控制技术存在一个问题,即由于隐马尔可夫模型没有明确的参数来对音素持续时间进行适当的建模,因此很难控制合成语音的音素持续时间。为了克服这个问题,我们使用多重回归隐半马尔可夫模型(MRHSMM),它具有显式的状态持续时间分布来控制电话持续时间。我们表明,持续时间控制是重要的风格控制的合成语音的主观测试的结果。我们还比较了所提出的技术与另一种基于模型插值的控制技术。
This paper presents a technique for controlling intuitively the degree or intensity of speaking styles and emotional expressions of synthetic speech. The conventional style control technique based on multiple regression HMM (MRHMM) has a problem that it is difficult to control phone duration of synthetic speech because HMM has no explicit parameter which models phone duration appropriately. To overcome this problem, we use multiple regression hidden semi-Markov model (MRHSMM) which has explicit state duration distributions to control phone duration. We show that the duration control is important for style control of synthetic speech from the results of subjective tests. We also compare the proposed technique with another control technique based on model interpolation.