Multi-lingual duration modeling

Multi-lingual duration modeling
复制标题

多语言持续时间建模

DOI:
10.21437/eurospeech.1997-669
复制
发表时间:
1997
期刊:
5th International Conference on Spoken Language Processing (ICSLP 1998)
影响因子:
--
通讯作者:
M. Tanenblatt
M. Tanenblatt
中科院分区:
--
文献类型:
--
作者:
J. V. Santen;Chilin Shih;Bernd Möbius;E. Tzoukermann;M. Tanenblatt

文献摘要

被引文献

相似文献

控制文本到语音合成系统中的时序很复杂,因为有许多上下文因素会影响时序;此外,哪些因素很重要以及它们的具体影响因语言而异。我们在这里描述了一种独立于语言的持续时间控制方法。在运行时,独立于语言的计时模块访问特定于语言的表。这些表指定了特征空间的哪些子类(即上下文和电话标识的所有组合)在特定意义上是同质的,即相同的因素对子类中的情况具有相似的影响。在子类中,持续时间通过简单的算术模型建模,例如乘法、加法或更普遍的乘积和模型。探索性统计方法(有监督)和参数估计技术(无监督)用于
Controlling timing in text-to-speech synthesis systems is complicated, because there are many contextual factors that affect timing; moreover, which factors matter and what their precise effects are varies among languages. We describe here a language-independent approach for duration control. At run time, a language-independent timing module accesses languagespecific tables. These tables specify which sub-classes of the feature space (i.e., all combinations of context and phone identity) are homogeneous in the specific sense that the same factors have similar effects on the cases in a sub-class. Within a sub-class, durations are modeled by simple arithmetic models such as multiplicative, additive, or – more generally – sums-ofproducts models. Exploratory statistical methods (supervised) and parameter estimation techniques (unsupervised) are used for