Singing voice analysis and editing based on mutually dependent F0 estimation and source separation

Singing voice analysis and editing based on mutually dependent F0 estimation and source separation
复制标题

DOI:
10.1109/icassp.2015.7178034
复制
发表时间:
2015-04
期刊:
2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Yukara Ikemiya;Kazuyoshi Yoshii;Katsutoshi Itoyama
Yukara Ikemiya;Kazuyoshi Yoshii;Katsutoshi Itoyama
中科院分区:
其他
文献类型:
--
作者:
Yukara Ikemiya;Kazuyoshi Yoshii;Katsutoshi Itoyama

文献摘要

被引文献

相似文献

本文提出了一种新的框架,通过有效利用声乐基本频率(F0)估计和歌声分离这两项任务的相互依赖性来改进这两项任务。歌唱声分离的典型方法是从目标音乐信号中估计歌唱声F0轮廓,然后通过使用仅通过歌唱声F0和泛音的谐波分量的时频掩模来提取歌唱声。相反,如果仅演唱声可以从目标信号中准确地提取,则认为人声F0估计变得更容易。这种相互依赖性在大多数传统研究中几乎没有得到关注。为了克服这一限制,我们的框架交替这两个任务,同时在另一个中使用每个任务的结果。更具体地说,我们首先提取歌声使用鲁棒主成分分析(RPCA)。然后,通过基于次谐波求和(SHS)的F0显着性谱图找到最佳路径,从分离的歌声中估计F0轮廓。这使我们能够提高演唱声分离相结合的时频掩模的基础上RPCA与基于谐波结构的掩模。当我们使用所提出的技术直接编辑流行音乐音频信号中的声乐F0时,得到的实验结果表明,它显着提高了声乐F0估计和歌声分离。
This paper presents a novel framework that improves both vocal fundamental frequency (F0) estimation and singing voice separation by making effective use of the mutual dependency of those two tasks. A typical approach to singing voice separation is to estimate the vocal F0 contour from a target music signal and then extract the singing voice by using a time-frequency mask that passes only the harmonic components of the vocal F0s and overtones. Vocal F0 estimation, on the contrary, is considered to become easier if only the singing voice can be extracted accurately from the target signal. Such mutual dependency has scarcely been focused on in most conventional studies. To overcome this limitation, our framework alternates those two tasks while using the results of each in the other. More specifically, we first extract the singing voice by using robust principal component analysis (RPCA). The F0 contour is then estimated from the separated singing voice by finding the optimal path over a F0-saliency spectrogram based on subharmonic summation (SHS). This enables us to improve singing voice separation by combining a time-frequency mask based on RPCA with a mask based on harmonic structures. Experimental results obtained when we used the proposed technique to directly edit vocal F0s in popular-music audio signals showed that it significantly improved both vocal F0 estimation and singing voice separation.