Data-driven solo voice enhancement for jazz music retrieval

Data-driven solo voice enhancement for jazz music retrieval
复制标题

DOI:
10.1109/icassp.2017.7952145
复制
发表时间:
2017-03
期刊:
2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
S. Balke;C. Dittmar;J. Abeßer;Meinard Müller
S. Balke;C. Dittmar;J. Abeßer;Meinard Müller
中科院分区:
其他
文献类型:
--
作者:
S. Balke;C. Dittmar;J. Abeßer;Meinard Müller

文献摘要

被引文献

相似文献

检索音乐录音中的短单音查询是音乐信息检索(MIR)中一个具有挑战性的研究问题。在爵士乐中,给定独奏转录,一项检索任务是在音乐集中找到相应的(可能是复调的)录音。许多传统系统通过首先从录音中提取主要的 F0 轨迹,然后将提取的轨迹量化为音高,最后将所得的音高序列与单音查询进行比较来完成此类检索任务。在本文中,我们介绍了一种数据驱动的方法,避免了传统方法中涉及的艰难决策:给定完整音乐录音的时频(TF)表示和独奏转录的 TF 表示对,我们使用基于 DNN 的方法来学习将“复调”TF 表示转换为“单音”TF 表示的映射。这种变换可以被视为一种独奏语音增强。我们在爵士乐独奏检索场景中评估我们的方法,并将其与最先进的主要旋律提取方法进行比较。
Retrieving short monophonic queries in music recordings is a challenging research problem in Music Information Retrieval (MIR). In jazz music, given a solo transcription, one retrieval task is to find the corresponding (potentially polyphonic) recording in a music collection. Many conventional systems approach such retrieval tasks by first extracting the predominant F0-trajectory from the recording, then quantizing the extracted trajectory to musical pitches and finally comparing the resulting pitch sequence to the monophonic query. In this paper, we introduce a data-driven approach that avoids the hard decisions involved in conventional approaches: Given pairs of time-frequency (TF) representations of full music recordings and TF representations of solo transcriptions, we use a DNN-based approach to learn a mapping for transforming a “polyphonic” TF representation into a “monophonic” TF representation. This transform can be considered as a kind of solo voice enhancement. We evaluate our approach within a jazz solo retrieval scenario and compare it to a state-of-the-art method for predominant melody extraction.