Parallel-Data-Free Many-to-Many Voice Conversion Based on DNN Integrated with Eigenspace Using a Non-Parallel Speech Corpus

Parallel-Data-Free Many-to-Many Voice Conversion Based on DNN Integrated with Eigenspace Using a Non-Parallel Speech Corpus
复制标题

基于非并行语音语料库与特征空间集成的 DNN 的并行无数据多对多语音转换

DOI:
10.21437/interspeech.2017-961
复制
发表时间:
2017
期刊:
--
影响因子:
--
通讯作者:
N. Minematsu
N. Minematsu
中科院分区:
--
文献类型:
--
作者:
Tetsuya Hashimoto;Hidetsugu Uchida;D. Saito;N. Minematsu

文献摘要

被引文献

相似文献

提出了一种新的无并行数据多对多语音转换方法。由于1对1转换的灵活性较差,研究人员主要关注多对多转换,其中说话人身份通常使用说话人空间基来表示。在这种情况下,同一句子的话语必须从许多说话者那里收集。本研究旨在克服这一限制,实现无并行数据的多对多转换。这是通过使用非并行语音语料库将深度神经网络(dnn)与特征空间相结合而实现的。在我们之前的研究中,使用DNN实现多对多转换,DNN的训练由EVGMM转换辅助。通过用非并行语料库构造特征空间等价地实现EVGMM函数,实现了期望的转换。这里的一个关键技术是在不给定源和目标说话人之间的并行数据的情况下估计协方差项。实验表明,该系统的客观评价分数与使用并行数据训练的基线系统相当。
This paper proposes a novel approach to parallel-data-free and many-to-many voice conversion (VC). As 1-to-1 conversion has less flexibility, researchers focus on many-to-many conversion, where speaker identity is often represented using speaker space bases. In this case, utterances of the same sentences have to be collected from many speakers. This study aims at overcoming this constraint to realize a parallel-data-free and many-to-many conversion. This is made possible by integrating deep neural networks (DNNs) with eigenspace using a nonparallel speech corpus. In our previous study, many-to-many conversion was implemented using DNN, whose training was assisted by EVGMM conversion. By realizing the function of EVGMM equivalently by constructing eigenspace with a nonparalell speech corpus, the desired conversion is made possible. A key technique here is to estimate covariance terms without given parallel data between source and target speakers. Experiments show that objective assessment scores are comparable to those of the baseline system trained with parallel data.