The Problem of Limited Inter-rater Agreement in Modelling Music Similarity

The Problem of Limited Inter-rater Agreement in Modelling Music Similarity
复制标题

音乐相似性建模中评估者间一致性有限的问题

DOI:
--
复制
发表时间:
2016
影响因子:
1.1
通讯作者:
Thomas Grill
Thomas Grill
中科院分区:
计算机科学4区
文献类型:
--
作者:
A. Flexer;Thomas Grill

文献摘要

参考文献

被引文献

相似文献

音乐信息检索(MIR)的核心目标之一是量化音乐片段之间或内部的相似性。这些定量关系应该反映人类对音乐相似性的感知,然而这是高度主观的,评分者之间的一致性很低。不幸的是,这一主要问题迄今为止在MIR中很少受到关注。由于计算模型超出人类一致性的水平是没有意义的,因此这些水平的评分者间一致性为任何算法方法提供了一个自然的上限。我们将说明这一基本问题,在MIR系统的评估使用两个典型的应用场景的结果:(i)音乐作品之间的音乐相似性建模;(ii)音乐作品内的音乐结构分析。对于这两个应用程序,我们推导出性能的上限,这是由于有限的评分员之间的协议。我们比较这些上限的国家的最先进的MIR系统的性能,并显示如何阻止更好的MIR系统的进一步发展的上限。
One of the central goals of Music Information Retrieval (MIR) is the quantification of similarity between or within pieces of music. These quantitative relations should mirror the human perception of music similarity, which is however highly subjective with low inter-rater agreement. Unfortunately this principal problem has been given little attention in MIR so far. Since it is not meaningful to have computational models that go beyond the level of human agreement, these levels of inter-rater agreement present a natural upper bound for any algorithmic approach. We will illustrate this fundamental problem in the evaluation of MIR systems using results from two typical application scenarios: (i) modelling of music similarity between pieces of music; (ii) music structure analysis within pieces of music. For both applications, we derive upper bounds of performance which are due to the limited inter-rater agreement. We compare these upper bounds to the performance of state-of-the-art MIR systems and show how the upper bounds prevent further progress in developing better MIR systems.
DOI: 10.1109/tmm.2014.2310701
发表时间: 2014-08-01
影响因子: 7.3
作者:
Serra, Joan;Mueller, Meinard;Arcos, Josep Ll
通讯作者: Arcos, Josep Ll