Statistical Voice Conversion with WaveNet-Based Waveform Generation

Statistical Voice Conversion with WaveNet-Based Waveform Generation
复制标题

DOI:
10.21437/interspeech.2017-986
复制
发表时间:
2017-08
影响因子:
4.2
通讯作者:
Kazuhiro Kobayashi;Tomoki Hayashi;Akira Tamamori;T. Toda
Kazuhiro Kobayashi;Tomoki Hayashi;Akira Tamamori;T. Toda
中科院分区:
工程技术3区
文献类型:
--
作者:
Kazuhiro Kobayashi;Tomoki Hayashi;Akira Tamamori;T. Toda

文献摘要

被引文献

相似文献

本文提出了一种基于 WaveNet 的波形生成的统计语音转换 (VC) 技术。基于高斯混合模型(GMM)的VC可以将源说话人的说话人身份转换为目标说话人的说话人身份。然而,在传统的声编码过程中,F 0 提取误差、参数化误差以及转换后的特征轨迹的过度平滑效应等多种因素会导致语音波形的建模误差,这通常会导致转换后的语音的音质下降。为了解决这个问题,我们将基于 WaveNet 声码器的直接波形生成技术应用于 VC。在所提出的方法中,首先,基于GMM将源说话人的声学特征转换为目标说话人的声学特征。然后,基于以转换后的声学特征为条件的 WaveNet 声码器生成转换后的语音的波形样本。在本文中,为了研究转换后的语音波形的建模精度,我们比较了基于 WaveNet 声码器训练和合成的几种声学特征。实验结果证实,与传统的 VC 技术相比,所提出的 VC 技术在扬声器个性方面实现了更高的转换精度,并且音质相当。
This paper presents a statistical voice conversion (VC) technique with the WaveNet-based waveform generation. VC based on a Gaussian mixture model (GMM) makes it possible to convert the speaker identity of a source speaker into that of a target speaker. However, in the conventional vocoding process, various factors such as F 0 extraction errors, parameterization errors and over-smoothing effects of converted feature trajectory cause the modeling errors of the speech waveform, which usually bring about sound quality degradation of the converted voice. To address this issue, we apply a direct waveform generation technique based on a WaveNet vocoder to VC. In the proposed method, first, the acoustic features of the source speaker are converted into those of the target speaker based on the GMM. Then, the waveform samples of the converted voice are generated based on the WaveNet vocoder conditioned on the converted acoustic features. In this paper, to investigate the modeling accuracies of the converted speech waveform, we compare several types of the acoustic features for training and synthesizing based on the WaveNet vocoder. The experimental re-sults confirmed that the proposed VC technique achieves higher conversion accuracy on speaker individuality with comparable sound quality compared to the conventional VC technique.