Glottal Spectral Separation for Speech Synthesis

Glottal Spectral Separation for Speech Synthesis
复制标题

用于语音合成的声门频谱分离

DOI:
10.1109/jstsp.2014.2307274
复制
发表时间:
2014
影响因子:
7.5
通讯作者:
Cabral J
Cabral J
中科院分区:
工程技术1区
文献类型:
--
作者:
Cabral J

文献摘要

相似文献

本文提出了一种分离声门源和声道成分的分析方法,称为声门谱分离(GSS)。该方法可以使用声门声源模型产生高质量的合成语音。在语音技术应用中常用的源滤波器模型中,假设源是频谱平坦的激励信号,并且声道滤波器可以由语音的频谱包络表示。虽然该模型可以产生高质量的语音,但它对语音转换具有限制,因为它不允许控制与语音质量相关的声门参数。使用更好地表示声门源和声道滤波器的语音模型的主要问题是,用于分离这些分量的当前分析方法不足以鲁棒到产生与使用基于语音的频谱包络的模型相同的语音质量。提出的GSS方法是克服这个问题的尝试,由以下三个步骤组成。最初,从语音信号估计声门源信号。然后,将语音频谱除以声门源信号的频谱包络,以便从语音信号中去除声门源效应。最后,声道传递函数是通过计算得到的信号的频谱包络。在这项工作中,声门源信号表示使用Liljencrants-Fant模型(LF模型)。我们在这里提出的实验表明,基于GSS的分析-合成技术可以产生的语音相媲美的高品质的声码器,是基于频谱包络表示。然而,它也允许控制语音质量,即通过修改声门参数将模态语音转换为呼吸和紧张。
This paper proposes an analysis method to separate the glottal source and vocal tract components of speech that is called Glottal Spectral Separation (GSS). This method can produce high-quality synthetic speech using an acoustic glottal source model. In the source-filter models commonly used in speech technology applications it is assumed the source is a spectrally flat excitation signal and the vocal tract filter can be represented by the spectral envelope of speech. Although this model can produce high-quality speech, it has limitations for voice transformation because it does not allow control over glottal parameters which are correlated with voice quality. The main problem with using a speech model that better represents the glottal source and the vocal tract filter is that current analysis methods for separating these components are not robust enough to produce the same speech quality as using a model based on the spectral envelope of speech. The proposed GSS method is an attempt to overcome this problem, and consists of the following three steps. Initially, the glottal source signal is estimated from the speech signal. Then, the speech spectrum is divided by the spectral envelope of the glottal source signal in order to remove the glottal source effects from the speech signal. Finally, the vocal tract transfer function is obtained by computing the spectral envelope of the resulting signal. In this work, the glottal source signal is represented using the Liljencrants-Fant model (LF-model). The experiments we present here show that the analysis-synthesis technique based on GSS can produce speech comparable to that of a high-quality vocoder that is based on the spectral envelope representation. However, it also permit control over voice qualities, namely to transform a modal voice into breathy and tense, by modifying the glottal parameters.