Singing Voice Enhancement in Monaural Music Signals Based on Two-stage Harmonic/Percussive Sound Separation on Multiple Resolution Spectrograms

Singing Voice Enhancement in Monaural Music Signals Based on Two-stage Harmonic/Percussive Sound Separation on Multiple Resolution Spectrograms
复制标题

DOI:
10.1109/taslp.2013.2287052
复制
发表时间:
2014
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Hideyuki Tachibana;Nobutaka Ono;S. Sagayama
Hideyuki Tachibana;Nobutaka Ono;S. Sagayama
中科院分区:
其他
文献类型:
--
作者:
Hideyuki Tachibana;Nobutaka Ono;S. Sagayama

文献摘要

相似文献

本文提出了一种针对单耳音乐音频信号的歌唱语音增强技术,这是一个非常具有挑战性的问题。最近提出了许多歌唱声音增强技术。然而,我们的方法是基于与这些现有方法完全不同的想法。我们以歌声的波动为研究对象,考虑利用两种不同分辨率的声谱图进行检测,一种声谱图具有高时间分辨率和低频率分辨率,另一种声谱图具有高频率分辨率和低时间分辨率。在这两个谱图上,波动分量的形状是完全不同的。基于这个想法,我们提出了一种歌唱声音增强技术,我们称之为两级谐波/打击声分离(HPSS)。在本文中,我们描述了两阶段HPSS的细节,并评估了该方法的性能。实验结果表明,作为一种常用的任务准则,SDR比现有的方法提高了约4 dB,这是一个相当高的水平。此外,我们还评估了该方法作为音乐旋律估计预处理的性能。实验结果表明,我们的歌声增强技术大大提高了简单的音高估计技术的性能。这些结果证明了所提方法的有效性。
We propose a novel singing voice enhancement technique for monaural music audio signals, which is a quite challenging problem. Many singing voice enhancement techniques have been proposed recently. However, our approach is based on a quite different idea from these existing methods. We focused on the fluctuation of a singing voice and considered to detect it by exploiting two differently resolved spectrograms, one has rich temporal resolution and poor frequency resolution, while the other has rich frequency resolution and poor temporal resolution. On such two spectrograms, the shapes of fluctuating components are quite different. Based on this idea, we propose a singing voice enhancement technique that we call two-stage harmonic/percussive sound separation (HPSS). In this paper, we describe the details of two-stage HPSS and evaluate the performance of the method. The experimental results show that SDR, a commonly-used criterion on the task, was improved by around 4 dB, which is a considerably higher level than existing methods. In addition, we also evaluated the performance of the method as a preprocessing for melody estimation in music. The experimental results show that our singing voice enhancement technique considerably improved the performance of a simple pitch estimation technique. These results prove the effectiveness of the proposed method.