Dilated Convolution with Dilated GRU for Music Source Separation

Dilated Convolution with Dilated GRU for Music Source Separation
复制标题

DOI:
10.24963/ijcai.2019/655
复制
发表时间:
2019-06
期刊:
--
影响因子:
--
通讯作者:
Jen-Yu Liu;Yi-Hsuan Yang
Jen-Yu Liu;Yi-Hsuan Yang
中科院分区:
其他
文献类型:
--
作者:
Jen-Yu Liu;Yi-Hsuan Yang

文献摘要

被引文献

相似文献

WaveNET中使用的堆叠膨胀卷积已被证明是产生高质量音频的有效方法。通过在褶积层中用膨胀来代替合并/跨步,它们可以保留高分辨率信息,并且仍然可以到达较远的位置。产生高分辨率的预测在音乐源分离中也是至关重要的,其目标是在保持分离的声音质量的同时分离不同的声源。因此,在本文中,我们使用堆叠膨胀卷积作为音乐信源分离的主干。虽然堆叠的膨胀卷积比标准卷积可以到达更广泛的背景,但它们的有效接收范围仍然是固定的,并且对于复杂的音乐音频信号可能不够宽。为了在更远的位置获得更多信息,我们建议将扩张卷积与称为扩张GRU的改进GRU相结合以形成块。扩展的GRU从之前的k步接收信息,而不是固定k步的前一步。这种修改允许GRU单元以较少的重复步骤到达位置,并且运行得更快,因为它可以部分并行执行。我们表明,在分离人声和伴奏方面,所提出的模型具有与现有技术同样好或更好的效果。
Stacked dilated convolutions used in Wavenet have been shown effective for generating high-quality audios. By replacing pooling/striding with dilation in convolution layers, they can preserve high-resolution information and still reach distant locations. Producing high-resolution predictions is also crucial in music source separation, whose goal is to separate different sound sources while maintain the quality of the separated sounds. Therefore, in this paper, we use stacked dilated convolutions as the backbone for music source separation. Although stacked dilated convolutions can reach wider context than standard convolutions do, their effective receptive fields are still fixed and might not be wide enough for complex music audio signals. To reach even further information at remote locations, we propose to combine a dilated convolution with a modified GRU called Dilated GRU to form a block. A Dilated GRU receives information from k-step before instead of the previous step for a fixed k. This modification allows a GRU unit to reach a location with fewer recurrent steps and run faster because it can execute in parallel partially. We show that the proposed model with a stack of such blocks performs equally well or better than the state-of-the-art for separating both vocals and accompaniment.