DEEP CLUSTERING AND CONVENTIONAL NETWORKS FOR MUSIC SEPARATION: STRONGER TOGETHER.

DEEP CLUSTERING AND CONVENTIONAL NETWORKS FOR MUSIC SEPARATION: STRONGER TOGETHER.
复制标题

DOI:
10.1109/icassp.2017.7952118
复制
发表时间:
2017-03
期刊:
Proceedings of the ... IEEE International Conference on Acoustics, Speech, and Signal Processing. ICASSP (Conference)
影响因子:
--
通讯作者:
Mesgarani N
Mesgarani N
中科院分区:
其他
文献类型:
--
作者:
Luo Y;Chen Z;Hershey JR;Le Roux J;Mesgarani N

文献摘要

被引文献

相似文献

深度聚类是第一种处理具有相同类型和任意数量的多个源的一般音频分离场景的方法,在与说话人无关的语音分离任务中表现出色。然而,人们对其在其他具有挑战性的情况(例如音乐源分离)中的有效性知之甚少。与直接估计源信号的传统网络相反,深度聚类为每个时频仓生成嵌入,并通过在嵌入空间中对仓进行聚类来分离源。我们表明,在匹配和不匹配的条件下,深度聚类在歌声分离任务上都优于传统网络,尽管传统网络具有最佳信号近似的端到端训练的优势,大概是因为其更灵活的目标可以产生更好的正则化。由于深度聚类和传统网络架构的优势似乎是互补的,因此我们探索将它们组合到通过类似于多任务学习的方法进行训练的单个混合网络中。值得注意的是,该组合的性能明显优于其中任何一个组件。
Deep clustering is the first method to handle general audio separation scenarios with multiple sources of the same type and an arbitrary number of sources, performing impressively in speaker-independent speech separation tasks. However, little is known about its effectiveness in other challenging situations such as music source separation. Contrary to conventional networks that directly estimate the source signals, deep clustering generates an embedding for each time-frequency bin, and separates sources by clustering the bins in the embedding space. We show that deep clustering outperforms conventional networks on a singing voice separation task, in both matched and mismatched conditions, even though conventional networks have the advantage of end-to-end training for best signal approximation, presumably because its more flexible objective engenders better regularization. Since the strengths of deep clustering and conventional network architectures appear complementary, we explore combining them in a single hybrid network trained via an approach akin to multi-task learning. Remarkably, the combination significantly outperforms either of its components.