Separation of Singing Voice Using Nonnegative Matrix Partial Co-Factorization for Singer Identification

Separation of Singing Voice Using Nonnegative Matrix Partial Co-Factorization for Singer Identification
复制标题

使用非负矩阵部分协分解分离歌声以进行歌手识别

DOI:
10.1109/taslp.2015.2396681
复制
发表时间:
2015-04
影响因子:
5.4
通讯作者:
Liu, Guizhong
Liu, Guizhong
中科院分区:
计算机科学2区
文献类型:
--
作者:
Hu, Ying;Liu, Guizhong

文献摘要

参考文献

相似文献

为了提高歌手识别的性能,我们提出了一种将单声道录音的歌声与音乐伴奏分开的系统。我们的系统由两个关键阶段组成。第一阶段利用非负矩阵部分协因式分解(NMPCF),这是一种结合歌声和纯伴奏先验知识的联合矩阵分解,将混合信号分离为歌声部分和伴奏部分。第二阶段,根据第一阶段分离得到的歌声,首先估计歌声的音高,然后可以区分歌声的谐波分量。对于一个帧来说,所区分的谐波分量被认为是可靠的,而其他频率分量则被认为是不可靠的,因此频谱是不完整的。利用这些谐波分量,可以通过缺失特征法、频谱重建来重建完整的歌声频谱,从而获得歌声更加干净的精细信号。实验结果表明,从音源分离的角度来看,歌声细化相对于使用NMPCF的歌声分离可以进一步提高ΔSNR,而从歌手识别的角度来看,NMPCF分离的歌声比细化后的歌声更合适。
In order to improve the performance of singer identification, we propose a system to separate singing voice from music accompaniment for monaural recordings. Our system consists of two key stages. The first stage exploits the nonnegative matrix partial co-factorization (NMPCF), which is a joint matrix decomposition integrating prior knowledge of singing voice and pure accompaniment to separate the mixture signal into singing voice portion and accompaniment portion. In the second stage, based on the separated singing voice obtained by the first stage, the pitches of singing voice are first estimated and then the harmonic components of singing voice can be distinguished. For a frame, the distinguished harmonic components are regarded as reliable while other frequency components unreliable, thus the spectrum is incomplete. With those harmonic components, the complete spectrums of singing voice can be reconstructed by a missing feature method, spectrum reconstruction, obtaining a refined signal with more clean singing voice. Experimental results demonstrate that, from the point view of source separation, the singing voice refinement can further improve ΔSNR in contrast with the singing voice separation using NMPCF, while for the point view of singer identification, the singing voice separated by NMPCF is more appropriate than the refined singing voice.
DOI: 10.1109/tasl.2009.2034186
发表时间: 2010-03
期刊: IEEE Transactions on Audio, Speech, and Language Processing
影响因子: --
作者:
E. Vincent;N. Bertin;R. Badeau
通讯作者: E. Vincent;N. Bertin;R. Badeau
DOI: 10.1109/aspaa.2005.1540227
发表时间: 2005-11
期刊: IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, 2005.
影响因子: --
作者:
Anssi Klapuri
通讯作者: Anssi Klapuri
DOI: 10.1109/tasl.2010.2042124
发表时间: 2010-11
期刊: IEEE Transactions on Audio, Speech, and Language Processing
影响因子: --
作者:
V. Rao;P. Rao
通讯作者: V. Rao;P. Rao
DOI: --
发表时间: 2008
期刊: --
影响因子: --
作者:
T. Virtanen;A. Mesaros;M. Ryynänen
通讯作者: T. Virtanen;A. Mesaros;M. Ryynänen
DOI: 10.1016/j.specom.2004.03.007
发表时间: 2004-09
期刊: Speech Commun.
影响因子: --
作者:
B. Raj;M. Seltzer;R. Stern
通讯作者: B. Raj;M. Seltzer;R. Stern