Analysis of minimum distances in high-dimensional musical spaces

Analysis of minimum distances in high-dimensional musical spaces
复制标题

DOI:
10.1109/tasl.2008.925883
复制
发表时间:
2008-07-01
影响因子:
--
通讯作者:
Slaney, Malcolm
Slaney, Malcolm
中科院分区:
其他
文献类型:
--
作者:
Casey, Michael;Rhodes, Christophe;Slaney, Malcolm

文献摘要

被引文献

相似文献

我们提出了一种自动方法,用于测量基于内容的音乐相似性,增强当前的音乐搜索引擎和推荐系统。以前的许多跟踪相似性的方法都需要蛮力,数据库中所有音频功能之间的配对处理,因此对于大型集合而言是不切实际的。但是,在一个与互联网连接的世界中,用户可以访问数百万次音乐曲目,效率至关重要。我们的方法使用通过分析确定的距离阈值从未标记的音频数据和接近Neigbor检索中提取的功能来求解一系列检索任务。这些任务需要时间功能 - 用于文本检索的摇摇欲坠技术。为了衡量相似性,我们在查询和目标轨道之间计算成对的音频瓦,距离阈值低于距离阈值。对于每个数据库,衬垫间距离的分布不同。因此,我们对带状疱疹之间最小距离的分布和一种估计最佳检索性能距离阈值的方法进行了分析。该方法与局部敏感的哈希(LSH)兼容,以比使用详尽的距离计算的速度更快地实现了几个数量级。我们评估了我们提出的方法在三个对比的音乐相似性任务上的性能:检索错误的录音(指纹),检索不同艺术家(封面歌曲)所执行的相同作品以及编辑和取样的查询曲目版本由混音艺术家(混音)。我们的方法在前两个任务中实现了接近完美的性能,在第三任任务中以70%的召回率达到75%的精度。每个任务均在包括450万音频瓦的测试数据库上执行。
We propose an automatic method for measuring content-based music similarity, enhancing the current generation of music search engines and recommender systems. Many previous approaches to track similarity require brute-force, pair-wise processing between all audio features in a database and therefore are not practical for large collections. However, in an Internet-connected world, where users have access to millions of musical tracks, efficiency is crucial. Our approach uses features extracted from unlabeled audio data and near-neigbor retrieval using a distance threshold, determined by analysis, to solve a range of retrieval tasks. The tasks require temporal features-analogous to the technique of shingling used for text retrieval. To measure similarity, we count pairs of audio shingles, between a query and target track, that are below a distance threshold. The distribution of between-shingle distances is different for each database; therefore, we present an analysis of the distribution of minimum distances between shingles and a method for estimating a distance threshold for optimal retrieval performance. The method is compatible with locality-sensitive hashing (LSH)-allowing implementation with retrieval times several orders of magnitude faster than those using exhaustive distance computations. We evaluate the performance of our proposed method on three contrasting music similarity tasks: retrieval of mis-attributed recordings (fingerprint), retrieval of the same work performed by different artists (cover songs), and retrieval of edited and sampled versions of a query track by remix artists (remixes). Our method achieves near-perfect performance in the first two tasks and 75% precision at 70% recall in the third task. Each task was performed on a test database comprising 4.5 million audio shingles.