Environmental Sound Recognition With Time-Frequency Audio Features

Environmental Sound Recognition With Time-Frequency Audio Features
复制标题

DOI:
10.1109/tasl.2009.2017438
复制
发表时间:
2009-08-01
影响因子:
--
通讯作者:
Kuo, C. -C. Jay
Kuo, C. -C. Jay
中科院分区:
其他
文献类型:
--
作者:
Chu, Selina;Narayanan, Shrikanth;Kuo, C. -C. Jay

文献摘要

被引文献

相似文献

本文认为,识别环境声音的理解的场景或上下文周围的音频传感器的任务。已经提出了各种特征用于音频识别,包括描述音频频谱形状的流行的Mel频率倒谱系数(MFCC)。环境声音(例如,昆虫的啁啾声及雨的声音,其通常为具有宽平坦频谱的类噪声)可包含强时间域签名。然而,只有少数的时域特征已被开发来表征这样的不同的音频信号之前。在这里,我们执行的音频环境表征的经验特征分析,并建议使用匹配追踪(MP)算法来获得有效的时频特征。基于MP的方法利用原子字典进行特征选择,从而产生灵活、直观和物理上可解释的特征集。采用基于MP的特征来补充MFCC特征,以获得更高的环境声音识别精度。进行了大量的实验,以证明这些联合功能的有效性,非结构化的环境声音分类,包括听力测试,以研究人类的识别能力。我们的识别系统已被证明可以产生与人类听众相当的性能。
The paper considers the task of recognizing environmental sounds for the understanding of a scene or context surrounding an audio sensor. A variety of features have been proposed for audio recognition, including the popular Mel-frequency cepstral coefficients (MFCCs) which describe the audio spectral shape. Environmental sounds, such as chirpings of insects and sounds of rain which are typically noise-like with a broad flat spectrum, may include strong temporal domain signatures. However, only few temporal-domain features have been developed to characterize such diverse audio signals previously. Here, we perform an empirical feature analysis for audio environment characterization and propose to use the matching pursuit (MP) algorithm to obtain effective time-frequency features. The MP-based method utilizes a dictionary of atoms for feature selection, resulting in a flexible, intuitive and physically interpretable set of features. The MP-based feature is adopted to supplement the MFCC features to yield higher recognition accuracy for environmental sounds. Extensive experiments are conducted to demonstrate the effectiveness of these joint features for unstructured environmental sound classification, including listening tests to study human recognition capabilities. Our recognition system has shown to produce comparable performance as human listeners.