Learning Environmental Sounds with Multi-scale Convolutional Neural Network

Learning Environmental Sounds with Multi-scale Convolutional Neural Network
复制标题

DOI:
10.1109/ijcnn.2018.8489641
复制
发表时间:
2018-03
期刊:
2018 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Boqing Zhu;Changjian Wang;Feng Liu;Jin Lei;Zengquan Lu;Yuxing Peng
Boqing Zhu;Changjian Wang;Feng Liu;Jin Lei;Zengquan Lu;Yuxing Peng
中科院分区:
其他
文献类型:
--
作者:
Boqing Zhu;Changjian Wang;Feng Liu;Jin Lei;Zengquan Lu;Yuxing Peng

文献摘要

被引文献

相似文献

深度学习极大地提高了声音识别的性能。然而,直接从原始波形中学习声学模型仍然具有挑战性。目前基于波形的模型一般采用时域卷积层提取特征。单一大小的滤波器提取的特征不足以建立音频的判别表示。在本文中,我们提出了多尺度卷积运算,通过提高频率分辨率和学习跨所有频率区域的滤波器来获得更好的音频表示。为了在单一模型中利用基于波形的特征和基于谱图的特征,我们引入了两阶段方法来融合不同的特征。最后,我们提出了一种基于多尺度卷积运算和两相方法的新型端到端网络WaveMsNet。在环境声音分类数据集ESC-10和ESC-50上,WaveMsNet的分类准确率分别达到93.75%和79.10%,较之前的方法有显著提高。
Deep learning has dramatically improved the performance of sounds recognition. However, learning acoustic models directly from the raw waveform is still challenging. Current waveform-based models generally use time-domain convolutional layers to extract features. The features extracted by single size filters are insufficient for building discriminative representation of audios. In this paper, we propose multi-scale convolution operation, which can get better audio representation by improving the frequency resolution and learning filters cross all frequency area. For leveraging the waveform-based features and spectrogram-based features in a single model, we introduce twophase method to fuse the different features. Finally, we propose a novel end-to-end network called WaveMsNet based on the multi-scale convolution operation and two-phase method. On the environmental sounds classification datasets ESC-10 and ESC-50, the classification accuracies of our WaveMsNet achieve 93.75% and 79.10% respectively, which improve significantly from the previous methods.