Combining pitch-based inference and non-negative spectrogram factorization in separating vocals from polyphonic music

Combining pitch-based inference and non-negative spectrogram factorization in separating vocals from polyphonic music
复制标题

DOI:
--
复制
发表时间:
2008
期刊:
--
影响因子:
--
通讯作者:
T. Virtanen;A. Mesaros;M. Ryynänen
T. Virtanen;A. Mesaros;M. Ryynänen
中科院分区:
其他
文献类型:
--
作者:
T. Virtanen;A. Mesaros;M. Ryynänen

文献摘要

被引文献

相似文献

本文提出了一种新的算法,用于分离人声从复调音乐伴奏。基于音高估计,该方法首先创建一个二进制掩码,指示在幅度谱图中的时频段,其中存在声乐信号的谐波内容。其次,非负矩阵分解(NMF)应用于非声乐段的声谱图,以学习模型的伴奏。NMF预测人声片段中的噪声量,这允许分离人声和噪声,即使它们在时间和频率上重叠。与基于正弦建模的参考算法相比,商业和合成声学材料的仿真分别显示出1.3 dB和1.8 dB的平均改善,并且分离的人声的感知质量也明显改善。该方法还在对齐分离的人声和文本歌词方面进行了测试,结果比参考方法更好。
This paper proposes a novel algorithm for separating vocals from polyphonic music accompaniment. Based on pitch estimation, the method first creates a binary mask indicating timefrequency segments in the magnitude spectrogram where harmonic content of the vocal signal is present. Second, nonnegative matrix factorization (NMF) is applied on the non-vocal segments of the spectrogram in order to learn a model for the accompaniment. NMF predicts the amount of noise in the vocal segments, which allows separating vocals and noise even when they overlap in time and frequency. Simulations with commercial and synthesized acoustic material show an average improvement of 1.3 dB and 1.8 dB, respectively, in comparison with a reference algorithm based on sinusoidal modeling, and also the perceptual quality of the separated vocals is clearly improved. The method was also tested in aligning separated vocals and textual lyrics, where it produced better results than the reference method.