Factorized context modelling for Text-to-Speech synthesis

Factorized context modelling for Text-to-Speech synthesis
复制标题

用于文本到语音合成的分解上下文建模

DOI:
10.1109/icassp.2013.6639192
复制
发表时间:
2013
期刊:
2013 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
通讯作者:
Simon King
Simon King
中科院分区:
--
文献类型:
--
作者:
Heng Lu;Simon King

文献摘要

参考文献

被引文献

相似文献

由于语音单元是如此依赖于上下文,大量的语言上下文特征通常用于基于HMM的文本到语音(TTS)语音合成系统,通过上下文相关的模型。由于不可能为每个上下文训练单独的模型,因此使用决策树来发现应该建模的最重要的特征组合。决策树的任务是非常困难的-从上下文特征空间的一个非常小的观察部分推广到其余部分-它们有一个主要的弱点:它们不能直接利用阶乘属性:它们每次基于一个特征细分模型空间。我们提出了一个基于动态贝叶斯网络(DBN)的混合记忆马尔可夫模型(MMMM),提供因式分解的上下文空间。听力测试的结果作为证据,该模型成功地学习了这个空间的阶乘性质。
Because speech units are so context-dependent, a large number of linguistic context features are generally used by HMM-based Text-to-Speech (TTS) speech synthesis systems, via context-dependent models. Since it is impossible to train separate models for every context, decision trees are used to discover the most important combinations of features that should be modelled. The task of the decision tree is very hard - to generalize from a very small observed part of the context feature space to the rest - and they have a major weakness: they cannot directly take advantage of factorial properties: they subdivide the model space based on one feature at a time. We propose a Dynamic Bayesian Network (DBN) based Mixed Memory Markov Model (MMMM) to provide factorization of the context space. The results of a listening test are provided as evidence that the model successfully learns the factorial nature of this space.
DOI: 10.1109/icassp.2000.861820
发表时间: 2000-06
期刊: 2000 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.00CH37100)
影响因子: --
作者:
K. Tokuda;Takayoshi Yoshimura;T. Masuko;Takao Kobayashi;T. Kitamura
通讯作者: K. Tokuda;Takayoshi Yoshimura;T. Masuko;Takao Kobayashi;T. Kitamura