Feature fusion for music detection

Feature fusion for music detection
复制标题

用于音乐检测的特征融合

DOI:
10.21437/eurospeech.1999-485
复制
发表时间:
1999
期刊:
--
影响因子:
--
通讯作者:
H. Lloyd
H. Lloyd
中科院分区:
--
文献类型:
--
作者:
Eluned S. Parris;M. Carey;H. Lloyd

文献摘要

被引文献

相似文献

近年来,音乐、语音和噪音之间的自动识别作为一个研究课题变得越来越重要。需要将音频分类为音乐或语音等类别是多媒体文档检索问题的重要部分。本文扩展了作者先前在音乐和语音识别中基于倒频、幅度、过零和音高比较静态和过渡特征的性能的工作。本文描述了两种方法来组合这些特性以提高整体性能。第一种方法为每个特征类型使用单独的GMM分类器,并融合分类器的输出。第二种方法是在用GMM对数据建模之前,将不同的特征组合成单个向量。与使用单一类型的特征所获得的结果相比,使用这两种方法可以观察到显著的性能改进。在使用17小时的测试材料进行的10秒测试中,最佳系统的错误率为0.3%。通过减少测试文件的长度,仅用两秒钟的数据就实现了小于1%的错误率,从而保持了性能。
Automatic discrimination between music, speech and noise has grown in importance as a research topic over recent years. The need to classify audio into categories such as music or speech is an important part of the multimedia document retrieval problem. This paper extends work previously carried out by the authors which compared performance of static and transitional features based on cepstra, amplitude, zerocrossings and pitch for music and speech discrimination. Two approaches are described to combine the features to improve overall performance. The first approach uses separate GMM classifiers for each feature type and fuses the outputs of the classifiers. The second approach combines different features into a single vector prior to modelling the data with a GMM. Significant improvements in performance have been observed using both approaches over the results achieved by a single type of feature. An equal error rate of 0.3% is achieved for the best system on ten second tests using seventeen hours of test material. The performance is maintained as the length of test file is reduced with an equal error rate of less than 1% being achieved with only two seconds of data.