Raw Multi-Channel Audio Source Separation using Multi- Resolution Convolutional Auto-Encoders

Raw Multi-Channel Audio Source Separation using Multi- Resolution Convolutional Auto-Encoders
复制标题

DOI:
10.23919/eusipco.2018.8553571
复制
发表时间:
2018-03
期刊:
2018 26th European Signal Processing Conference (EUSIPCO)
影响因子:
--
通讯作者:
Emad M. Grais;D. Ward;Mark D. Plumbley
Emad M. Grais;D. Ward;Mark D. Plumbley
中科院分区:
其他
文献类型:
--
作者:
Emad M. Grais;D. Ward;Mark D. Plumbley

文献摘要

被引文献

相似文献

监督多声道音频源分离需要从混合信号中提取有用的频谱,时间和空间特征。因此,许多现有系统的成功很大程度上取决于用于训练的特征的选择。在这项工作中,我们引入了一种新颖的多通道,多分辨率卷积自编码器神经网络,该神经网络处理原始时域信号,以确定适当的多分辨率特征,用于将歌唱声音从立体声音乐中分离出来。实验结果表明,该方法可以实现多声道音源分离,不需要手工制作特征,也不需要任何预处理或后处理。
Supervised multi-channel audio source separation requires extracting useful spectral, temporal, and spatial features from the mixed signals. the success of many existing systems is therefore largely dependent on the choice of features used for training. In this work, we introduce a novel multi-channel, multiresolution convolutional auto-encoder neural network that works on raw time-domain signals to determine appropriate multiresolution features for separating the singing-voice from stereo music. Our experimental results show that the proposed method can achieve multi-channel audio source separation without the need for hand-crafted features or any pre- or post-processing.