Merging data driven and rule based prosodic models for unit selection TTS

Merging data driven and rule based prosodic models for unit selection TTS
复制标题

合并数据驱动和基于规则的韵律模型以进行单元选择 TTS

DOI:
--
复制
发表时间:
2004
期刊:
--
影响因子:
--
通讯作者:
M. Aylett
M. Aylett
中科院分区:
--
文献类型:
--
作者:
M. Aylett

文献摘要

被引文献

相似文献

数据驱动的模型受到数据稀疏性的影响,很难泛化。基于规则的模型的缺点是过于规定性,而且对单位选择数据库的内容不敏感。更复杂的是,任何一个话语的可接受韵律空间都很大。然而,在某些情况下,特定说话者的韵律模式可能是非常均匀的,例如,用于读出邮政编码的韵律模式。本文描述了一种在有限域内探索和分析韵律空间的方法,以及一种将基于规则的简单韵律模型与一组数据驱动的迷你韵律模型合并的方法。在有和没有迷你模型的情况下进行了邮政编码合成的听力测试,取得了令人满意的结果。这种方法可以有效地应用于从数字金额到个人姓名的各种域名。
Data driven models suffer from data sparsity and can be difficult to generalise. Rule based models suffer from being over prescriptive and insensitive to the contents of the unit selection database. To further complicate matters the space of acceptable prosody for any one utterance is large. However in some cases prosodic patterns for a particular speaker can be very homogeneous, for example the prosodic pattern used to read out a zip code. In this paper we describe a method for exploring and analysing the prosodic space within a limited domain, and a method for merging a simple rule based prosodic model with a set of data driven mini prosodic models. A listening test was carried out on the synthesis of zip codes with and without the mini models with promising results. The approach could be applied effectively to domains varying from numerical amounts to personal names.