Few-Shot Musical Source Separation

Few-Shot Musical Source Separation
复制标题

少镜头音乐源分离

DOI:
10.1109/icassp43922.2022.9747536
复制
发表时间:
2022
期刊:
Speech and Signal Processing (ICASSP
影响因子:
--
通讯作者:
Pablo Bello, Juan
Pablo Bello, Juan
中科院分区:
--
文献类型:
--
作者:
Wang, Yu;Stoller, Daniel;Bittner, Rachel M.;Pablo Bello, Juan

文献摘要

参考文献

被引文献

相似文献

基于深度学习的音乐源分离方法通常限于模型训练的乐器类别,并且不能概括为分离看不见的乐器。为了解决这个问题,我们提出了一个少数镜头的音乐源分离范例。我们使用目标乐器的几个音频示例来调节通用的U-Net源分离模型。我们与U-Net一起训练了一个少拍条件编码器,将音频示例编码为条件向量,以通过特征线性调制(FiLM)配置U-Net。我们在MUSDB 18和MedleyDB数据集中对真实的音乐录音进行了训练模型的评估。我们表明,我们提出的几杆空调范例优于基线一热仪器类条件模型,看到和看不见的仪器。为了将我们的方法扩展到更广泛的现实世界场景,我们还尝试了不同的条件反射示例特征,包括来自不同录音的示例,多个来源或负面条件反射示例。
Deep learning-based approaches to musical source separation are often limited to the instrument classes that the models are trained on and do not generalize to separate unseen instruments. To address this, we propose a few-shot musical source separation paradigm. We condition a generic U-Net source separation model using few audio examples of the target instrument. We train a few-shot conditioning encoder jointly with the U-Net to encode the audio examples into a conditioning vector to configure the U-Net via feature-wise linear modulation (FiLM). We evaluate the trained models on real musical recordings in the MUSDB18 and MedleyDB datasets. We show that our proposed few-shot conditioning paradigm outperforms the base-line one-hot instrument-class conditioned model for both seen and unseen instruments. To extend the scope of our approach to a wider variety of real-world scenarios, we also experiment with different conditioning example characteristics, including examples from different recordings, with multiple sources, or negative conditioning examples.
DOI: 10.1109/icassp39728.2021.9413896
发表时间: 2020
期刊: ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者:
Woosung Choi;Minseok Kim;Jaehwa Chung;Soonyoung Jung
通讯作者: Soonyoung Jung
用于音乐源分离的类条件嵌入
DOI: 10.1109/icassp.2019.8683007
发表时间: 2018
期刊: ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者:
Prem Seetharaman;G. Wichern;Shrikant Venkataramani;Jonathan Le Roux
通讯作者: Jonathan Le Roux
任意声音的一次性条件音频过滤
DOI: --
发表时间: 2020
期刊: IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子: --
作者:
Beat Gfeller;Dominik Roblek;M. Tagliasacchi
通讯作者: M. Tagliasacchi
Conditioned-U-Net:在 U-Net 中引入控制机制以实现多源分离
DOI: --
发表时间: 2019
期刊: International Society for Music Information Retrieval Conference
影响因子: --
作者:
Gabriel Meseguer;Geoffroy Peeters
通讯作者: Geoffroy Peeters
MedleyDB:用于注释密集型 MIR 研究的多轨数据集
DOI: --
发表时间: 2014
期刊: 15th International Society for Music Information Retrieval Conference
影响因子: --
作者:
Bittner, R.
通讯作者: Bittner, R.