Many-to-Many and Completely Parallel-Data-Free Voice Conversion Based on Eigenspace DNN

Many-to-Many and Completely Parallel-Data-Free Voice Conversion Based on Eigenspace DNN
复制标题

基于特征空间DNN的多对多完全并行无数据语音转换

DOI:
10.1109/taslp.2018.2878949
复制
发表时间:
2019
期刊:
IEEE/ACM Transaction on Audio, Speech and Language Processing
影响因子:
--
通讯作者:
Nobuaki Minematsu
Nobuaki Minematsu
中科院分区:
--
文献类型:
--
作者:
Tetsuya Hashimoto;Daisuke Saito;Nobuaki Minematsu

文献摘要

参考文献

被引文献

相似文献

图像、文字、语音等媒体转换,通常需要大量并行数据来训练转换模型。近年来,不使用或使用少量并行数据训练模型的方法引起了研究者的注意。在多对多语音转换中,由于通常难以从每对说话者收集并行数据,因此期望不需要并行数据的转换模型。传统的多对多语音转换模型需要大量的预存并行数据来获取整个说话人空间的先验知识。然后,从一个任意的扬声器到另一个特定的模型可以通过调整一些模型参数来实现。虽然这些转换模型在自适应步骤中肯定不使用并行数据,但它们仍然使用并行数据进行先前的训练。在这项研究中,我们的目标是实现完全并行的数据和多对多的语音转换。所提出的方法使用特征语音高斯混合模型(EVGMM)和深度神经网络(DNN)。EVGMM是一种多对多转换模型,它通过分析高斯混合模型的均值向量来构建整个说话人空间(称为本征空间),并且在我们的方法中使用它将训练说话人的特征分解到其本征空间分量中。通过使用说话人特征和所获得的分量作为伪并行数据,训练多个DNN以实现它们之间的转换。有了这些DNN,任何目标说话者的特征都可以用分量的加权和来表示。应该指出的是,我们的提案的所有过程都不需要任何并行数据。一个关键技术是估计EVGMM的协方差项没有平行的数据。实验表明,所提出的方法,不使用并行数据的个性得分是足够的可比性与并行数据训练的基线系统。
Media conversion of image, text, speech, etc., generally requires a large amount of parallel data for training a conversion model. Recently, methods for training the model using no or a small amount of parallel data draw researchers' attention. In many-to-many voice conversion, since it is often hard to collect parallel data from every pair of speakers, the conversion models requiring no parallel data are desired. Conventional many-to-many voice conversion models required a large amount of prestored parallel data to acquire prior knowledge of the entire speaker space. Then, a specific model from an arbitrary speaker to another can be realized by adapting a few model parameters. Although these conversion models certainly do not use parallel data in an adaptation step, they still use parallel data for prior training. In this study, we aim at realizing completely parallel-data-free and many-to-many voice conversion. The proposed method uses both Eigenvoice Gaussian mixture models (EVGMM) and Deep neural network (DNN). EVGMM is a many-to-many conversion model that constructs the entire speaker space (called eigenspace) by analyzing mean vectors of Gaussian mixture models and it is used in our method to decompose training speakers' features into their eigenspace components. By using the speaker features and the obtained components as pseudo parallel data, multiple DNNs are trained to realize conversion between them. With these DNNs, features of any target speaker can be represented by a weighted sum of the components. It should be noted that all the processes of our proposal do not require any parallel data. A key technique is to estimate covariance terms of EVGMM with no parallel data. Experiments indicate that individuality scores of the proposed method using no parallel data are comparable enough to those of a baseline system trained with parallel data.
基于非并行语音语料库与特征空间集成的 DNN 的并行无数据多对多语音转换
DOI: 10.21437/interspeech.2017-961
发表时间: 2017
期刊: --
影响因子: --
作者:
Tetsuya Hashimoto;Hidetsugu Uchida;D. Saito;N. Minematsu
通讯作者: N. Minematsu
使用动态核偏最小二乘回归对非并行数据集进行语音转换
DOI: 10.21437/interspeech.2013-103
发表时间: 2013
期刊: 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
影响因子: --
作者:
Hanna Silén;J. Nurminen;E. Helander;M. Gabbouj
通讯作者: M. Gabbouj
基于深度神经网络构建的说话人空间基的任意说话人转换
DOI: 10.1109/apsipa.2016.7820831
发表时间: 2016
期刊: 2016 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA)
影响因子: --
作者:
Tetsuya Hashimoto;D. Saito;N. Minematsu
通讯作者: N. Minematsu