Exploring Music Contents

Exploring Music Contents
复制标题

探索音乐内容

DOI:
10.1007/978-3-642-23126-1_10
复制
发表时间:
2011
期刊:
--
影响因子:
--
通讯作者:
Barthet M
Barthet M
中科院分区:
--
文献类型:
--
作者:
Barthet M

文献摘要

被引文献

相似文献

我们提出了两个语音/音乐的歧视方法,使用音色模型和测量他们的表现上3小时长的数据库广播播客从BBC。在第一种方法中,使用中值滤波对用自动音色识别(ATR)模型获得的机器估计分类进行后处理。分类系统(LSF/K-means)使用两个不同的分类级别进行训练,一个是高级(语音,音乐),另一个是低级(男性和女性语音,古典,爵士,摇滚和流行)。第二种方法结合自动结构分割和音色识别(ASS/ATR)。ASS使用HMM和软K均值算法评估特征分布(MFCC,RMS)之间的相似性。这两种方法进行了评估,在语义(相对正确的重叠RCO),和时间(边界检索F-措施)的水平。ASS/ATR方法获得了最好的结果(平均RCO为94.5%,边界F-测量为50.1%)。这些性能相比,得到了有利的基于SVM的技术提供了一个很好的基准的最先进的。
We propose two speech/music discrimination methods using timbre models and measure their performances on a 3 hour long database of radio podcasts from the BBC. In the first method, the machine estimated classifications obtained with an automatic timbre recognition (ATR) model are post-processed using median filtering. The classification system (LSF/K-means) was trained using two different taxonomic levels, a high-level one (speech, music), and a lower-level one (male and female speech, classical, jazz, rock & pop). The second method combines automatic structural segmentation and timbre recognition (ASS/ATR). The ASS evaluates the similarity between feature distributions (MFCC, RMS) using HMM and soft K-means algorithms. Both methods were evaluated at a semantic (relative correct overlap RCO), and temporal (boundary retrieval F-measure) levels. The ASS/ATR method obtained the best results (average RCO of 94.5% and boundary F-measure of 50.1%). These performances were favourably compared with that obtained by a SVM-based technique providing a good benchmark of the state of the art.