A Globally Regularized Joint Neural Architecture for Music Classification

A Globally Regularized Joint Neural Architecture for Music Classification
复制标题

用于音乐分类的全局正则化联合神经架构

DOI:
10.1109/access.2020.3043142
复制
发表时间:
2020-01-01
期刊:
影响因子:
3.9
通讯作者:
Abid, Fazeel
Abid, Fazeel
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ashraf, Mohsin;Geng, Guohua;Abid, Fazeel

文献摘要

被引文献

相似文献

音乐分类是音乐信息检索(MIR)在组织大量音乐收藏中的重要应用。以可靠的准确性对不同音乐进行分类的任务被认为是具有挑战性的。这些任务中的大多数采用手工特征工程来构建分类器,但无法识别音乐的原始特征。使用卷积神经网络(CNN)和递归神经网络(RNN)的神经网络的几种组合已经被许多研究人员考虑。然而,人们已经注意到,CNN和RNN的联合架构由于批量归一化而存在一些问题,这导致了低准确性和更多的训练时间。为了解决这些问题,提出了基于CNN和RNN的混合模型的全局层正则化(GLR)技术,使用Mel谱图来评估训练和准确性。我们的实验使用很少的超参数,通过分别实现87.79%和68.87%的适度准确度,提高了GTZAN和Free Music Achieve(FMA)数据集的性能。从经验上讲,我们提出的模型的时空域功能和全局层正则化技术的优势,以实现可靠的准确性相比,其他国家的最先进的作品。
Music classification is an essential application of Music Information Retrieval (MIR) in organizing extensive collections of music. The tasks to classify different music with reliable accuracy observed to be challenging. Most of these tasks employ handcrafted feature engineering to build a classifier, yet unable to identify the original characteristics of music. Several combinations of neural networks using convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have been in consideration of many researchers. However, it has been noticed that the joint architecture of CNN and RNN suffers some problems due to batch normalization, which causes low accuracy and more training time. To handle these issues, the Global Layer Regularization (GLR) technique is proposed on the hybrid model of CNN and RNN using Mel-spectrograms for the evaluation of training and accuracy. Our experiments, with few hyper-parameters, improve performance on GTZAN and Free Music Achieve (FMA) datasets by achieving modest accuracy of 87.79% and 68.87% respectively. Empirically, our proposed model takes the advantages of spatiotemporal domain features and the global layer regularization technique to accomplish reliable accuracy as compared to the other state of art works.