Automatic segmentation combining an HMM-based approach and spectral boundary correction

Automatic segmentation combining an HMM-based approach and spectral boundary correction
复制标题

结合基于 HMM 的方法和光谱边界校正的自动分割

DOI:
--
复制
发表时间:
2002
期刊:
Interspeech
影响因子:
--
通讯作者:
Alistair Conkie
Alistair Conkie
中科院分区:
--
文献类型:
--
作者:
Yeon;Alistair Conkie

文献摘要

被引文献

相似文献

目前,AT&T实验室的Natural Voices多语言TTS系统通过大规模语音语料库产生高质量的合成语音[1]。在这种系统的开发中,自动分割构成了主要的组成技术。语音合成中的自动切分是基于隐马尔可夫模型(HMM)的。尽管基于HMM的方法是最自动和可靠的,但仍然存在一些限制,例如手工标记的transmittance和HMM对齐标签之间的不匹配,这可能导致合成语音中的不连续性,或者在HMM初始化中需要手工标记的引导数据。本文介绍了一种新的自动分割方法,其目的是最大限度地减少人为干预,并实现更高的分段质量的合成语音单元拼接合成,通过结合传统的基于HMM的方法和频谱边界校正。偏好测试表明,所提出的方法是有效的,减少合成语音中的不连续性。
Currently, AT&T Labs’ Natural Voices multilingual TTS system produces high-quality synthetic speech with a large-scale speech corpus [1]. In the development of such systems, automatic segmentation constitutes a major component technology. The prevalent approach for automatic segmentation in speech synthesis is Hidden Markov Model (HMM) - based. Even though an HMM-based approach is the most automatic and reliable, there are still several limitations, such as mismatches between hand-labeled transcriptions and HMM alignment labels which can lead to discontinuities in the synthetic speech, or the need for hand-labeled bootstrap data in HMM initialization. This paper introduces a new approach to automatic segmentation which aims both to minimize human intervention and to achieve a higher segmental quality of synthetic speech in unit-concatenative speech synthesis, by combining a conventional HMM-based approach and spectral boundary correction. A preference test demonstrates the proposed method is effective in reducing discontinuities in synthetic speech.