A practical singing voice detection system based on GRU-RNN
A practical singing voice detection system based on GRU-RNN
复制标题
一种基于GRU-RNN的实用歌声检测系统
DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
W. Li
中科院分区:
文献类型:
--
作者:
Z. Chen;X. Zhang;J. Deng;J. Li;Y. Jiang;W. Li
In this paper, we present a practical three-step approach for singing voice detection based on a gated recurrent unit (GRU) recurrent neural network (RNN) and the proposed method achieves comparable results to state-of-the-art method. We combine four classic features—namely Mel-frequency Cepstral Coefficients (MFCC), Mel-filter Bank, Linear Predictive Cepstral Coefficients (LPCC), and Chroma. Then, the mixed signal is first preprocessed by singing voice separation (SVS) with the Deep U-Net Convolutional Networks. Long short-term memory (LSTM) and GRU are both proposed to solve the gradient vanish problem in RNN. In our experiments, we set the block duration as 120 ms and 720 ms respectively, and we get comparable or better results than results from state-of-the-art methods, while results on Jamendo are not as good as those from RWC-Pop.