Factorized context modelling for Text-to-Speech synthesis
Factorized context modelling for Text-to-Speech synthesis
复制标题
用于文本到语音合成的分解上下文建模
DOI:
10.1109/icassp.2013.6639192
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
Simon King
中科院分区:
文献类型:
--
作者:
Heng Lu;Simon King
Because speech units are so context-dependent, a large number of linguistic context features are generally used by HMM-based Text-to-Speech (TTS) speech synthesis systems, via context-dependent models. Since it is impossible to train separate models for every context, decision trees are used to discover the most important combinations of features that should be modelled. The task of the decision tree is very hard - to generalize from a very small observed part of the context feature space to the rest - and they have a major weakness: they cannot directly take advantage of factorial properties: they subdivide the model space based on one feature at a time. We propose a Dynamic Bayesian Network (DBN) based Mixed Memory Markov Model (MMMM) to provide factorization of the context space. The results of a listening test are provided as evidence that the model successfully learns the factorial nature of this space.
DOI:
10.1109/icassp.2000.861820
发表时间:
2000-06
期刊:
2000 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.00CH37100)
影响因子:
--
作者:
K. Tokuda;Takayoshi Yoshimura;T. Masuko;Takao Kobayashi;T. Kitamura
通讯作者:
K. Tokuda;Takayoshi Yoshimura;T. Masuko;Takao Kobayashi;T. Kitamura