Non-Parallel Voice Conversion Using Variational Autoencoders Conditioned by Phonetic Posteriorgrams and D-Vectors

Non-Parallel Voice Conversion Using Variational Autoencoders Conditioned by Phonetic Posteriorgrams and D-Vectors
复制标题

DOI:
10.1109/icassp.2018.8461384
复制
发表时间:
2018-04
期刊:
2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Yuki Saito;Yusuke Ijima;Kyosuke Nishida;Shinnosuke Takamichi
Yuki Saito;Yusuke Ijima;Kyosuke Nishida;Shinnosuke Takamichi
中科院分区:
其他
文献类型:
--
作者:
Yuki Saito;Yusuke Ijima;Kyosuke Nishida;Shinnosuke Takamichi

文献摘要

被引文献

相似文献

本文提出了一种基于变分自编码器(VAEs)的非并行语音转换框架。尽管传统的基于vae的VC模型可以使用具有给定说话人表示的非并行语音语料库进行训练,但由于在vae的潜在变量中经常观察到过度正则化问题,转换后的语音内容往往会消失。为了克服这一问题,本文提出了一种基于语音识别的非并行语音识别方法,该方法不仅以说话人的表征为条件,而且以语音后图(ppg)表示语音的语音内容。由于语音内容是在训练过程中给出的,我们可以期望VC模型能够有效地学习与说话人无关的语音潜在特征。针对这一点,本文还将传统的基于虚拟语音识别的非并行语音识别扩展为多对多语音识别,可以将任意说话人的特征转换为另一个任意说话人的特征。我们研究了两种方法来估计用于训练VC模型的未包含在语音语料库中的说话人的说话人表示:1)采用传统的说话人编码,以及2)使用d-向量来表示说话人。实验结果表明:1)PPGs成功地提高了转换后语音的自然度和说话人相似度;2)说话人编码和d向量都可以用于基于vae的多对多非并行VC。
This paper proposes novel frameworks for non-parallel voice conversion (VC) using variational autoencoders (VAEs). Although conventional VAE-based VC models can be trained using non-parallel speech corpora with given speaker representations, phonetic contents of the converted speech tend to vanish because of an over-regularization issue often observed in latent variables of the VAEs. To overcome the issue, this paper proposes a VAE-based non-parallel VC conditioned by not only the speaker representations but also phonetic contents of speech represented as phonetic posteriorgrams (PPGs). Since the phonetic contents are given during the training, we can expect that the VC models effectively learn speaker-independent latent features of speech. Focusing on the point, this paper also extends the conventional VAE-based non-parallel VC to many-to-many VC that can convert arbitrary speakers' characteristics into another arbitrary speakers' ones. We investigate two methods to estimate speaker representations for speakers not included in speech corpora used for training VC models: 1) adapting conventional speaker codes, and 2) using d-vectors for the speaker representations. Experimental results demonstrate that 1) PPGs successfully improve both naturalness and speaker similarity of the converted speech, and 2) both speaker codes and d-vectors can be adopted to the VAE-based many-to-many non-parallel VC.