A practical singing voice detection system based on GRU-RNN

A practical singing voice detection system based on GRU-RNN
复制标题

一种基于GRU-RNN的实用歌声检测系统

DOI:
--
复制
发表时间:
2019
期刊:
Springer LNEE 568 (Lecture Notes in Electrical Engineering)
影响因子:
--
通讯作者:
W. Li
W. Li
中科院分区:
其他
文献类型:
--
作者:
Z. Chen;X. Zhang;J. Deng;J. Li;Y. Jiang;W. Li

文献摘要

相似文献

在本文中,我们提出了一个实用的三步方法的基础上门控递归单元(GRU)的递归神经网络(RNN)的歌声检测和所提出的方法取得了可比的结果,以国家的最先进的方法。我们结合了联合收割机的四个经典功能,即梅尔频率倒谱系数(MFCC),梅尔滤波器组,线性预测倒谱系数(LPCC),和色度。然后,混合信号首先通过深U网卷积网络的歌声分离(SVS)进行预处理。长短期记忆(LSTM)和GRU都被提出来解决RNN中的梯度消失问题。在我们的实验中,我们将块持续时间分别设置为120 ms和720 ms,我们得到了与最先进的方法相当或更好的结果,而Jamendo的结果不如RWC-Pop的结果。
In this paper, we present a practical three-step approach for singing voice detection based on a gated recurrent unit (GRU) recurrent neural network (RNN) and the proposed method achieves comparable results to state-of-the-art method. We combine four classic features—namely Mel-frequency Cepstral Coefficients (MFCC), Mel-filter Bank, Linear Predictive Cepstral Coefficients (LPCC), and Chroma. Then, the mixed signal is first preprocessed by singing voice separation (SVS) with the Deep U-Net Convolutional Networks. Long short-term memory (LSTM) and GRU are both proposed to solve the gradient vanish problem in RNN. In our experiments, we set the block duration as 120 ms and 720 ms respectively, and we get comparable or better results than results from state-of-the-art methods, while results on Jamendo are not as good as those from RWC-Pop.