A scale-free distribution of false positives for a large class of audio similarity measures

A scale-free distribution of false positives for a large class of audio similarity measures
复制标题

DOI:
10.1016/j.patcog.2007.04.012
复制
发表时间:
2008-01-01
影响因子:
8
通讯作者:
Pachet, Francois
Pachet, Francois
中科院分区:
计算机科学1区
文献类型:
--
作者:
Aucouturier, Jean-Julien;Pachet, Francois

文献摘要

被引文献

相似文献

音频模式识别的“帧袋”方法(13017)将信号建模为其局部频谱特征的长期统计分布,其原型实现是梅尔频率倒谱系数的高斯混合模型。这种方法是从音乐信号中提取高级描述的最主要范例,例如它们的乐器,流派或情绪,并且还可以用于计算歌曲之间的直接音色相似性。然而,作者最近的一项研究表明,这类算法在应用于音乐时往往会产生误报,无论查询如何,这些误报通常都是相同的歌曲。换句话说,在这样的模型中,存在着一些歌曲我们称之为中心,它们与很多歌曲无关地接近。本文报告了一些实验,使用大型音乐数据库上的实现,旨在更好地了解这种集线器歌曲的性质和原因。我们引入了两个措施的“hubness”,n-出现的数量和平均邻居角。我们发现,在典型的音乐数据库中,中心点是沿着无标度分布分布的:非中心点的歌曲非常普遍,而大的中心点非常罕见,但它们确实存在。此外,我们建立了枢纽不是给定建模策略的属性(即静态与动态,参数与非参数等)。而是倾向于在任何类型的模型中发生,然而仅对于具有给定量的“异质性”(待定义)的数据。这表明枢纽的存在可能是一个重要的现象,概括了音乐建模的具体问题,并表明了一类重要的模式识别算法的一般结构属性。(C)2007模式识别学会。由爱思唯尔有限公司出版。保留所有权利。
The "bag-of-frames" approach (13017) to audio pattern recognition models signals as the long-term statistical distribution of their local spectral features, a prototypical implementation of which being Gaussian Mixture Models of Mel-Frequency Cepstrum Coefficients. This approach is the most predominant paradigm to extract high-level descriptions from music signals, such as their instrument, genre or mood, and can also be used to compute direct timbre similarity between songs. However, a recent study by the authors shows that this class of algorithms when applied to music tends to create false positives which are mostly always the same songs regardless of the query. In other words, with such models, there exist songs-which we call hubs-which are irrelevantly close to very many songs. This paper reports on a number of experiments, using implementations on large music databases, aiming at better understanding the nature and causes of such hub songs. We introduce two measures of "hubness", the number of n-occurrences and the mean neighbor angle. We find that in typical music databases, hubs are distributed along a scale-free distribution: non-hub songs are extremely common, and large hubs are extremely rare-but they exist. Moreover, we establish that hubs are not a property of a given modelling strategy (i.e. static vs dynamic, parametric vs non-parametric, etc.) but rather tend to occur with any type of model, however only for data with a given amount of "heterogeneity" (to be defined). This suggests that the existence of hubs could be an important phenomenon which generalizes over the specific problem of music modelling, and indicates a general structural property of an important class of pattern recognition algorithms. (C) 2007 Pattern Recognition Society. Published by Elsevier Ltd. All rights reserved.