Robust Sound Event Classification Using Deep Neural Networks

Robust Sound Event Classification Using Deep Neural Networks
复制标题

使用深度神经网络进行稳健的声音事件分类

DOI:
10.1109/taslp.2015.2389618
复制
发表时间:
2015-03-01
影响因子:
5.4
通讯作者:
Xiao, Wei
Xiao, Wei
中科院分区:
计算机科学2区
文献类型:
--
作者:
McLoughlin, Ian;Zhang, Haomin;Xiao, Wei

文献摘要

被引文献

相似文献

计算机对声音事件的自动识别是自动监控、机器听觉和听觉场景理解等新兴应用的一个重要方面。机器学习以及人类听觉系统的计算模型的最新进展有助于这一日益流行的研究领域的进步。鲁棒的声音事件分类,即在现实世界嘈杂条件下识别声音的能力,是一项特别具有挑战性的任务。从语音识别域转换的分类方法,使用的功能,如梅尔频率倒谱系数,已被证明是执行合理的声音事件分类任务,虽然基于频谱或听觉图像分析技术,据报道,在噪声中实现上级性能。本文概述了一种声音事件分类框架,该框架使用支持向量机和深度神经网络分类器将听觉图像前端特征与基于声谱图图像的前端特征进行比较。在不同水平的破坏性噪音和几种系统增强下,对标准鲁棒分类任务的性能进行了评估,结果表明与当前最先进的分类技术相比非常好。
The automatic recognition of sound events by computers is an important aspect of emerging applications such as automated surveillance, machine hearing and auditory scene understanding. Recent advances in machine learning, as well as in computational models of the human auditory system, have contributed to advances in this increasingly popular research field. Robust sound event classification, the ability to recognise sounds under real-world noisy conditions, is an especially challenging task. Classification methods translated from the speech recognition domain, using features such as mel-frequency cepstral coefficients, have been shown to perform reasonably well for the sound event classification task, although spectrogram-based or auditory image analysis techniques reportedly achieve superior performance in noise. This paper outlines a sound event classification framework that compares auditory image front end features with spectrogram image-based front end features, using support vector machine and deep neural network classifiers. Performance is evaluated on a standard robust classification task in different levels of corrupting noise, and with several system enhancements, and shown to compare very well with current state-of-the-art classification techniques.