Noise-Robust Voice Conversion Based on Sparse Spectral Mapping Using Non-negative Matrix Factorization

Noise-Robust Voice Conversion Based on Sparse Spectral Mapping Using Non-negative Matrix Factorization
复制标题

DOI:
10.1587/transinf.e97.d.1411
复制
发表时间:
2014-06
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Ryo Aihara;R. Takashima;T. Takiguchi;Y. Ariki
Ryo Aihara;R. Takashima;T. Takiguchi;Y. Ariki
中科院分区:
其他
文献类型:
--
作者:
Ryo Aihara;R. Takashima;T. Takiguchi;Y. Ariki

文献摘要

相似文献

本文提出了一种基于语音稀疏表示的噪声环境语音转换(VC)技术。使用非负矩阵分解 (NMF) 的基于稀疏表示的 VC 用于不同说话者之间添加噪声的频谱转换。在我们之前基于样本的 VC 方法中,源样本和目标样本是从并行训练数据中提取的,具有由源和目标说话者说出的相同文本。输入源信号使用源样本及其权重来表示。然后,根据目标样本和与源样本相关的权重构建转换后的语音。然而,这种基于样本的方法需要保存所有训练样本(帧),并且需要大量计算时间才能获得源样本的权重。在本文中,我们提出了一个框架来训练源样本和目标样本的基础矩阵,以便它们具有共同的权重矩阵。通过使用基矩阵而不是样本,执行 VC 的计算时间比基于样本的方法要少。通过将其有效性(在使用添加噪声的语音数据的说话人转换实验中)与基于示例的方法和基于传统高斯混合模型(GMM)的方法的有效性进行比较,证实了该方法的有效性。关键词: 语音转换, 稀疏表示, 非负矩阵分解, 噪声鲁棒性
This paper presents a voice conversion (VC) technique for noisy environments based on a sparse representation of speech. Sparse representation-based VC using Non-negative matrix factorization (NMF) is employed for noise-added spectral conversion between different speakers. In our previous exemplar-based VC method, source exemplars and target exemplars are extracted from parallel training data, having the same texts uttered by the source and target speakers. The input source signal is represented using the source exemplars and their weights. Then, the converted speech is constructed from the target exemplars and the weights related to the source exemplars. However, this exemplar-based approach needs to hold all training exemplars (frames), and it requires high computation times to obtain the weights of the source exemplars. In this paper, we propose a framework to train the basis matrices of the source and target exemplars so that they have a common weight matrix. By using the basis matrices instead of the exemplars, the VC is performed with lower computation times than with the exemplar-based method. The effectiveness of this method was confirmed by comparing its effectiveness (in speaker conversion experiments using noise-added speech data) with that of an exemplar-based method and a conventional Gaussian mixture model (GMM)-based method. key words: voice conversion, sparse representation, non-negative matrix factorization, noise robustness