A style control technique for speech synthesis using multiple regression HSMM
A style control technique for speech synthesis using multiple regression HSMM
复制标题
DOI:
10.21437/interspeech.2006-388
复制
发表时间:
2006
期刊:
影响因子:
--
通讯作者:
Takashi Nose;J. Yamagishi;Takao Kobayashi
中科院分区:
文献类型:
--
作者:
Takashi Nose;J. Yamagishi;Takao Kobayashi
This paper presents a technique for controlling intuitively the degree or intensity of speaking styles and emotional expressions of synthetic speech. The conventional style control technique based on multiple regression HMM (MRHMM) has a problem that it is difficult to control phone duration of synthetic speech because HMM has no explicit parameter which models phone duration appropriately. To overcome this problem, we use multiple regression hidden semi-Markov model (MRHSMM) which has explicit state duration distributions to control phone duration. We show that the duration control is important for style control of synthetic speech from the results of subjective tests. We also compare the proposed technique with another control technique based on model interpolation.