MULTICHANNEL AUDIO CLASSIFICATION WITH NEURAL NETWORKS USING SCATTERING TRANSFORM

MULTICHANNEL AUDIO CLASSIFICATION WITH NEURAL NETWORKS USING SCATTERING TRANSFORM
复制标题

使用散射变换的神经网络多通道音频分类

DOI:
--
复制
发表时间:
2018
期刊:
--
影响因子:
--
通讯作者:
A. Amar
A. Amar
中科院分区:
--
文献类型:
--
作者:
S. Ezra;Y. Gershon;U. Levi;M. Palatin;A. Raveh;S. Sheer;Y. Doweck;A. Amar

文献摘要

被引文献

相似文献

本技术论文介绍了2018年声学场景分类挑战(DCASE 2018)任务5的方法。通过具有4个麦克风的阵列观察音频片段的序列。该任务是建议一个多通道处理,以分类的音频信号到9个预定义的类之一。所提出的方法将深度神经网络与散射变换相结合。每个音频片段首先由两层散射变换表示。将两层中的每一层的4个去噪变换组合在一起。每个融合层都由两个神经网络(NN)架构(RESNET和长短期记忆(LSTM)网络)并行处理,并具有一个联合的全连接层。
This technical paper presents an approach for the 2018 acoustic scene classification challenge (DCASE 2018) task 5. A sequence of audio segments are observed by an array with 4 microphones. The task is to suggest a multichannel processing to classify the audio signals to one of 9 pre-defined classes. The proposed approach combines a deep neural network with scattering transform. Each audio segment is first represented by two layers of scattering transform. The 4 denoised transforms of each of the two layers are combined together. Each of the fused layers are processed in parallel by two neural networks (NN) architectures, RESNET and long short-term memory (LSTM) network, with a joint fully connected layer.