An Improved Speech / Nonspeech Classification Based on Feature Combination for Audio Indexing

An Improved Speech / Nonspeech Classification Based on Feature Combination for Audio Indexing
复制标题

一种改进的基于音频索引特征组合的语音/非语音分类

DOI:
10.1587/transfun.e93.a.830
复制
发表时间:
2010
期刊:
IEICE Trans. Fundam. Electron. Commun. Comput. Sci.
影响因子:
--
通讯作者:
M. Hagiwara
M. Hagiwara
中科院分区:
--
文献类型:
--
作者:
Ji;Hyon;M. Hagiwara

文献摘要

被引文献

相似文献

在这封信中,我们提出了一种改进的语音/非语音分类方法来有效地对多媒体源进行分类。为了提高性能,我们引入了基于频谱持续时间分析的特征,并结合了最近提出的高过零率比(HZCRR)、低短时能量比(LSTER)和俯仰比(PR)等特征。根据我们对语音、音乐和环境声音的实验结果,与传统方法相比,所提出的方法获得了较高的分类结果。
In this letter, we propose an improved speech/nonspeech classification method to effectively classify a multimedia source. To improve performance, we introduce a feature based on spectral duration analysis, and combine recently proposed features such as high zero crossing rate ratio (HZCRR), low short time energy ratio (LSTER), and pitch ratio (PR). According to the results of our experiments on speech, music, and environmental sounds, the proposed method obtained high classification results when compared with conventional approaches.