Low-Rank Representation of Both Singing Voice and Music Accompaniment Via Learned Dictionaries

Low-Rank Representation of Both Singing Voice and Music Accompaniment Via Learned Dictionaries
复制标题

DOI:
--
复制
发表时间:
2013
期刊:
--
影响因子:
--
通讯作者:
Yi-Hsuan Yang
Yi-Hsuan Yang
中科院分区:
其他
文献类型:
--
作者:
Yi-Hsuan Yang

文献摘要

被引文献

相似文献

最近的研究工作表明,一首歌曲的幅度谱图可以被认为是一个低秩分量和稀疏分量的叠加,这似乎分别对应于歌曲的器乐部分和声乐部分。基于这种观察,人们可以将歌声从背景音乐中分离出来。然而,这种分离的质量可能是有限的,因为歌曲的声乐部分有时也可能是低秩的。因此,我们建议首先从一组干净的信号中学习声乐和乐器声音的子空间结构,然后根据学习到的子空间计算歌曲的声乐和乐器部分的低秩表示。具体来说,我们使用在线字典学习学习的子空间,并提出了一种新的算法称为多低秩表示(MLRR)的幅度谱图分解成两个低秩矩阵。我们的方法是灵活的,歌声和音乐伴奏的子空间都是从数据中学习。在MIR-1 K数据集上的实验结果表明,该方法提高了源失真比(SDR)和源干扰比(SIR),但对源伪影比(SAR)影响不大。
Recent research work has shown that the magnitude spectrogram of a song can be considered as a superposition of a low-rank component and a sparse component, which appear to correspond to the instrumental part and the vocal part of the song, respectively. Based on this observation, one can separate singing voice from the background music. However, the quality of such separation might be limited, because the vocal part of a song can sometimes be lowrank as well. Therefore, we propose to learn the subspace structures of vocal and instrumental sounds from a collection of clean signals first, and then compute the low-rank representations of both the vocal and instrumental parts of a song based on the learned subspaces. Specifically, we use online dictionary learning to learn the subspaces, and propose a new algorithm called multiple low-rank representation (MLRR) to decompose a magnitude spectrogram into two low-rank matrices. Our approach is flexible in that the subspaces of singing voice and music accompaniment are both learned from data. Evaluation on the MIR-1K dataset shows that the approach improves the source-to-distortion ratio (SDR) and the source-to-interference ratio (SIR), but not the source-to-artifact ratio (SAR).