A voice conversion method based on joint pitch and spectral envelope transformation

A voice conversion method based on joint pitch and spectral envelope transformation
复制标题

一种基于联合基音和谱包络变换的语音转换方法

DOI:
10.21437/interspeech.2004-452
复制
发表时间:
2004
期刊:
2009 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
通讯作者:
T. Chonavel
T. Chonavel
中科院分区:
--
文献类型:
--
作者:
T. En;O. Rosec;T. Chonavel

文献摘要

被引文献

相似文献

语音转换的研究大多集中在频谱变换上,而韵律特征的转换基本上是通过简单的基音周期线性变换来实现的。这些单独的转换导致不令人满意的语音转换质量,特别是当源说话人和目标说话人的说话风格不同时。在本文中,我们提出了一种能够联合转换基音和光谱包络信息的方法。要变换的参数是通过将缩放后的基音值与浊音帧的频谱包络参数和仅用于清音帧的频谱包络参数组合来获得的。使用高斯混合模型(GMM)对这些参数进行聚类。然后使用条件期望估计器确定变换函数。试验表明,该工艺取得了令人满意的基音转换效果。此外,它还使谱包络变换更加稳健。
Most of the research in Voice Conversion (VC) is devoted to spectral transformation while the conversion of prosodic features is essentially obtained through a simple linear transformation of pitch. These separate transformations lead to an unsatisfactory speech conversion quality, especially when the speaking styles of the source and target speakers are different. In this paper, we propose a method capable of jointly converting pitch and spectral envelope information. The parameters to be transformed are obtained by combining scaled pitch values with the spectral envelope parameters for the voiced frames and only spectral envelope parameters for the unvoiced ones. These parameters are clustered using a Gaussian Mixture Model (GMM). Then the transformation functions are determined using a conditional expectation estimator. Tests carried out show that, this process leads to a satisfactory pitch transformation. Moreover, it makes the spectral envelope transformation more robust.