A fast audio classification from MPEG coded data

A fast audio classification from MPEG coded data
复制标题

MPEG 编码数据的快速音频分类

DOI:
10.1109/icassp.1999.757473
复制
发表时间:
1999
期刊:
1999 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings. ICASSP99 (Cat. No.99CH36258)
影响因子:
--
通讯作者:
A. Kurematsu
A. Kurematsu
中科院分区:
--
文献类型:
--
作者:
Y. Nakajima;Yang Lu;M. Sugano;A. Yoneyama;H. Yanagihara;A. Kurematsu

文献摘要

被引文献

相似文献

音频信息分类在自动关键词识别等基于内容的音视频查询系统中是一项非常重要的任务。在本文中,我们描述了一种快速和准确的MPEG编码数据域上的音频数据分类方法。首先,无声段检测使用不同的记录条件下的鲁棒性的方法。然后利用子带能量的时间密度、带宽和中心频率将非静音段分为音乐、语音和掌声三类。为了尽可能地对各种音频源具有鲁棒性,我们使用贝叶斯判别函数对多变量高斯分布进行判别,而不是手动调整每个阈值。在实验中,每一秒的MPEG音频数据进行分类,约90%的音频和语音段已被成功检测。至于检测速度,需要不到MPEG音频解码处理能力的20%。
Audio information classification becomes a very important task for such purposes as automatic keyword spotting and other content-based audio-visual query systems. In this paper, we describe a fast and accurate audio data classification method on the MPEG coded data domain. Firstly silent segments are detected using a robust approach for different recording conditions. Then the non-silent segments are classified into three types, music, speech, and applause using temporal density, bandwidth and center frequency of subband energy. In order to be robust for a variety of audio sources as much as possible, we use Bayes discriminant function for multivariate Gaussian distribution instead of manually adjusting a threshold for each discriminator. In the experiment, every one-second of MPEG audio data is classified and about 90% of audio and speech segments have been successfully detected. As for the detection speed, less than 20% of MPEG audio decoding processing power is required.