Evolving Multi-Resolution Pooling CNN for Monaural Singing Voice Separation

Evolving Multi-Resolution Pooling CNN for Monaural Singing Voice Separation
复制标题

用于单声道歌声分离的不断发展的多分辨率池化 CNN

DOI:
10.1109/taslp.2021.3051331
复制
发表时间:
2020-08
期刊:
IEEE-ACM Transactions on Audio, Speech, and Language Processing (SCI 一区)
影响因子:
--
通讯作者:
Wenwu Wang
Wenwu Wang
中科院分区:
其他
文献类型:
--
作者:
Weitao Yuan;Bofei Dong;Shengbei Wang;Masashi Unoki;Wenwu Wang

文献摘要

参考文献

被引文献

相似文献

单声道歌唱声分离是一个具有挑战性的课题,并得到了广泛的研究。深度神经网络(DNN)是目前MSVS的最先进方法。然而,它们通常是手动设计的,这既耗时又容易出错。它们也是预定义的,因此不能使其结构适应训练数据。为了解决这些问题,我们首先为MSVS设计了一个多分辨率卷积神经网络(CNN),称为多分辨率池化CNN(MRP-CNN),它使用不同大小的池化算子来提取多分辨率特征。然后,我们引入了神经结构搜索(NAS),将MRP-CNN扩展到不断发展的MRP-CNN(E-MRP-CNN),使用遗传算法自动搜索有效的MRP-CNN结构,该遗传算法根据仅考虑分离性能的单个目标和考虑分离性能和模型复杂性的多个目标进行优化。使用多目标算法的E-MRP-CNN给出了一组帕累托最优解,每个解都提供了分离性能和模型复杂性之间的权衡。对MIR-1 K、DSD 100和MUSDB 18数据集的评估被用来证明E-MRP-CNN相对于几个最近的基线的优势。
Monaural singing voice separation (MSVS) is a challenging task and has been extensively studied. Deep neural networks (DNNs) are current state-of-the-art methods for MSVS. However, they are often designed manually, which is time-consuming and error-prone. They are also pre-defined, thus cannot adapt their structures to the training data. To address these issues, we first designed a multi-resolution convolutional neural network (CNN) for MSVS called multi-resolution pooling CNN (MRP-CNN), which uses various-sized pooling operators to extract multi-resolution features. We then introduced Neural Architecture Search (NAS) to extend the MRP-CNN to the evolving MRP-CNN (E-MRP-CNN) to automatically search for effective MRP-CNN structures using genetic algorithms optimized in terms of a single objective taking into account only separation performance and multiple objectives taking into account both separation performance and model complexity. The E-MRP-CNN using the multi-objective algorithm gives a set of Pareto-optimal solutions, each providing a trade-off between separation performance and model complexity. Evaluations on the MIR-1 K, DSD100, and MUSDB18 datasets were used to demonstrate the advantages of the E-MRP-CNN over several recent baselines.
DOI: --
发表时间: 2013
期刊: --
影响因子: --
作者:
Yi-Hsuan Yang
通讯作者: Yi-Hsuan Yang
DOI: --
发表时间: 2018-09
期刊: --
影响因子: --
作者:
Cheng-Zhi Anna Huang;Ashish Vaswani;Jakob Uszkoreit;Noam M. Shazeer;Ian Simon;Curtis Hawthorne;Andrew M. Dai;M. Hoffman;Monica Dinculescu;D. Eck
通讯作者: Cheng-Zhi Anna Huang;Ashish Vaswani;Jakob Uszkoreit;Noam M. Shazeer;Ian Simon;Curtis Hawthorne;Andrew M. Dai;M. Hoffman;Monica Dinculescu;D. Eck
DOI: 10.1109/icassp.2017.7952118
发表时间: 2017-03
期刊: Proceedings of the ... IEEE International Conference on Acoustics, Speech, and Signal Processing. ICASSP (Conference)
影响因子: --
作者:
Luo Y;Chen Z;Hershey JR;Le Roux J;Mesgarani N
通讯作者: Mesgarani N
DOI: 10.24963/ijcai.2019/655
发表时间: 2019-06
期刊: --
影响因子: --
作者:
Jen-Yu Liu;Yi-Hsuan Yang
通讯作者: Jen-Yu Liu;Yi-Hsuan Yang
DOI: --
发表时间: 2018-09
期刊: ArXiv
影响因子: --
作者:
Cheng-Zhi Anna Huang;Ashish Vaswani;Jakob Uszkoreit;Noam M. Shazeer;Curtis Hawthorne;Andrew M. Dai;M. Hoffman;D. Eck
通讯作者: Cheng-Zhi Anna Huang;Ashish Vaswani;Jakob Uszkoreit;Noam M. Shazeer;Curtis Hawthorne;Andrew M. Dai;M. Hoffman;D. Eck