Evaluating Prosodic Processing for Incremental Speech Synthesis

Evaluating Prosodic Processing for Incremental Speech Synthesis
复制标题

评估增量语音合成的韵律处理

DOI:
--
复制
发表时间:
2012
期刊:
Interspeech
影响因子:
--
通讯作者:
David Schlangen
David Schlangen
中科院分区:
--
文献类型:
--
作者:
Timo Baumann;David Schlangen

文献摘要

被引文献

相似文献

增量语音合成(iSS)接受输入并以连续的块产生输出,这些块仅一起产生完整的话语。因此,使用iSS的系统有能力在他们正在进行的时候调整他们的话语。然而,以少于可用的完整话语开始处理禁止了全局优化,导致潜在的次优解决方案。在本文中,我们提出了一种方法,用于递增的语音合成的符号预处理组件和评估的影响,不同的“前瞻”,即。e.关于话语其余部分的知识,关于韵律质量的知识。我们发现,高质量的增量输出,甚至可以实现一个前瞻小于一个短语,允许及时的系统反应。
Incremental speech synthesis (iSS) accepts input and produces output in consecutive chunks that only together result in a full utterance. Systems that use iSS thus have the ability to adapt their utterances while they are ongoing. However, starting to process with less than the full utterance available prohibits global optimization, leading to potentially suboptimal solutions. In this paper, we present a method for incrementalizing the symbolic pre-processing component of speech synthesis and assess the influence of varying “lookahead”, i. e. knowledge about the rest of the utterance, on prosodic quality. We found that high quality incremental output can be achieved even with a lookahead of less than one phrase, allowing for timely system reaction.