Voice conversion based on Non-negative Matrix Factorization in noisy environments

Voice conversion based on Non-negative Matrix Factorization in noisy environments
复制标题

DOI:
10.1109/sii.2013.6776630
复制
发表时间:
2013-12
期刊:
Proceedings of the 2013 IEEE/SICE International Symposium on System Integration
影响因子:
--
通讯作者:
Takao Fujii;Ryo Aihara;R. Takashima;T. Takiguchi;Y. Ariki
Takao Fujii;Ryo Aihara;R. Takashima;T. Takiguchi;Y. Ariki
中科院分区:
其他
文献类型:
--
作者:
Takao Fujii;Ryo Aihara;R. Takashima;T. Takiguchi;Y. Ariki

文献摘要

相似文献

提出了一种适用于噪声环境的语音转换技术。我们准备了平行样本(词典),由源样本和目标样本组成,这些样本具有源和目标说话者发出的相同文本。输入源信号被分解为源样本、从输入信号获得的噪声样本以及它们的权重(活动)。然后,通过计算目标样本和使用源样本计算的权重的线性组合来获得转换后的信号。在该方法中,还将基于高斯混合模型(GMM)的转换方法应用于稀疏编码生成的特征向量,以补偿源样本和目标样本的权重之间的不匹配。通过与传统方法的有效性比较,证实了该方法的有效性。
This paper presents a voice conversion (VC) technique for noisy environments. We prepared parallel exemplars (dictionary) that consist of the source and target exemplars, which have the same texts uttered by the source and target speakers. The input source signal is decomposed into the source exemplars, noise exemplars obtained from the input signal, and their weights (activities). Then, the converted signal is obtained by calculating the linear combination of the target exemplars and the weights which are calculated using the source exemplars. In the proposed method, a Gaussian Mixture Model (GMM) -based conversion method is also applied to the feature vectors generated by the sparse coding in order to compensate a mismatch between the weights of source and target exemplars. The effectiveness of this method was confirmed by comparing its effectiveness with that of a conventional method.