Extracting singing voice from music recordings by cascading audio decomposition techniques

Extracting singing voice from music recordings by cascading audio decomposition techniques
复制标题

通过级联音频分解技术从音乐录音中提取歌声

DOI:
10.1109/icassp.2015.7177945
复制
发表时间:
2015
期刊:
2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Meinard Müller
Meinard Müller
中科院分区:
--
文献类型:
--
作者:
Jonathan Driedger;Meinard Müller

文献摘要

参考文献

被引文献

相似文献

近年来,从音乐录音中提取歌声的问题引起了越来越多的研究兴趣。许多提出的分解技术都基于以下两种策略之一。第一种方法是利用对歌唱声音的特定特征的知识,将给定的音乐记录直接分解为一个用于歌唱声音的分量和一个用于伴奏的分量。第二种方法之后的程序将记录分解成一大组细粒度组件,这些组件被分类并随后重新组装,以产生所需的源估计。在本文中,我们提出了一种新的方法,结合了这两种策略的优点。我们首先以级联的方式应用不同的音频分解技术,将音乐记录分解成一组中级组件。这种分解可以很好地对歌唱声音的各种特征进行建模,但也足够粗略,可以保持成分的明确语义。这些属性允许我们直接从组件中重组歌声和伴奏。我们的客观和主观评价表明,该策略可以与最先进的歌唱语音分离算法相媲美,并获得了令人满意的结果。
The problem of extracting singing voice from music recordings has received increasing research interest in recent years. Many proposed decomposition techniques are based on one of the following two strategies. The first approach is to directly decompose a given music recording into one component for the singing voice and one for the accompaniment by exploiting knowledge about specific characteristics of singing voice. Procedures following the second approach disassemble the recording into a large set of fine-grained components, which are classified and reassembled afterwards to yield the desired source estimates. In this paper, we propose a novel approach that combines the strengths of both strategies. We first apply different audio decomposition techniques in a cascaded fashion to disassemble the music recording into a set of mid-level components. This decomposition is fine enough to model various characteristics of singing voice, but coarse enough to keep an explicit semantic meaning of the components. These properties allow us to directly reassemble the singing voice and the accompaniment from the components. Our objective and subjective evaluations show that this strategy can compete with state-of-the-art singing voice separation algorithms and yields perceptually appealing results.
DOI: 10.1007/978-0-387-30441-0
发表时间: 2008
期刊: --
影响因子: --
作者:
D. Havelock;桑野 園子;M. Vorländer
通讯作者: D. Havelock;桑野 園子;M. Vorländer
DOI: 10.1007/978-0-387-30441-0_23
发表时间: 2008
期刊: PACRIM. 2005 IEEE Pacific Rim Conference on Communications, Computers and signal Processing, 2005.
影响因子: --
作者:
Youngmoo E. Kim
通讯作者: Youngmoo E. Kim
DOI: 10.3390/ijms22126504
发表时间: 2021-06-17
影响因子: 5.6
作者:
Park PSU;Raynor WY;Sun Y;Werner TJ;Rajapakse CS;Alavi A
通讯作者: Alavi A