Arbitrary speaker conversion based on speaker space bases constructed by deep neural networks

Arbitrary speaker conversion based on speaker space bases constructed by deep neural networks
复制标题

基于深度神经网络构建的说话人空间基的任意说话人转换

DOI:
10.1109/apsipa.2016.7820831
复制
发表时间:
2016
期刊:
2016 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA)
影响因子:
--
通讯作者:
N. Minematsu
N. Minematsu
中科院分区:
--
文献类型:
--
作者:
Tetsuya Hashimoto;D. Saito;N. Minematsu

文献摘要

被引文献

相似文献

本文提出了一种构建基于深度神经网络 (DNN) 的语音转换 (VC) 系统的新方法,其中 DNN 与说话人特征空间集成。所提出的网络由多个 DNN 组成,每个 DNN 将输入特征转换为与特征空间基相对应的特征。这些 DNN 的训练是在 Eigenvoice GMM (EVGMM) 的帮助下实现的。使用一对多 VC 任务的实验评估表明,与 EVGMM 相比,该方法取得了更好的性能。
This paper proposes a novel approach to construct a Deep Neural Network (DNN) based voice conversion (VC) system, where DNNs are integrated with speaker eigenspace. The proposed network consists of multiple DNNs and each of them converts input features to features corresponding to a base of eigenspace. Training of these DNNs is achieved with the assistance of Eigenvoice GMM (EVGMM). Experimental evaluations using one-to-many VC tasks show that the proposed method achieved better performance compared with that of EVGMM.