Average-Voice-Based Speech Synthesis Using HSMM-Based Speaker Adaptation and Adaptive Training

Average-Voice-Based Speech Synthesis Using HSMM-Based Speaker Adaptation and Adaptive Training
复制标题

DOI:
10.1093/ietisy/e90-d.2.533
复制
发表时间:
2007-02
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
J. Yamagishi;Takao Kobayashi
J. Yamagishi;Takao Kobayashi
中科院分区:
其他
文献类型:
--
作者:
J. Yamagishi;Takao Kobayashi

文献摘要

被引文献

相似文献

在用于语音合成的说话人自适应中,期望转换语音特征和韵律特征,诸如F0和音素持续时间。为了在HMM框架内同时适应频谱、F0和音素持续时间,我们不仅需要变换对应于频谱和F0的状态输出分布,还需要变换对应于音素持续时间的持续时间分布。然而,由于原始HMM没有显式的持续时间分布,因此调整状态持续时间并不简单。因此,我们利用的隐半马尔可夫模型(HSMM),这是一个隐马尔可夫模型具有明确的状态持续时间分布的框架,我们应用基于HSMM的模型自适应算法,同时转换的状态输出和状态持续时间分布。此外,我们提出了一个基于HSMM的自适应训练算法,同时规范化的平均语音模型的状态输出和状态持续时间分布。我们将这些技术结合到我们的基于HSMM的语音合成系统,并从主观和客观评价测试的结果显示其有效性。
In speaker adaptation for speech synthesis, it is desirable to convert both voice characteristics and prosodic features such as F0 and phone duration. For simultaneous adaptation of spectrum, F0 and phone duration within the HMM framework, we need to transform not only the state output distributions corresponding to spectrum and F0 but also the duration distributions corresponding to phone duration. However, it is not straightforward to adapt the state duration because the original HMM does not have explicit duration distributions. Therefore, we utilize the framework of the hidden semi-Markov model (HSMM), which is an HMM having explicit state duration distributions, and we apply an HSMM-based model adaptation algorithm to simultaneously transform both the state output and state duration distributions. Furthermore, we propose an HSMM-based adaptive training algorithm to simultaneously normalize the state output and state duration distributions of the average voice model. We incorporate these techniques into our HSMM-based speech synthesis system, and show their effectiveness from the results of subjective and objective evaluation tests.