MIR_EVAL: A Transparent Implementation of Common MIR Metrics

MIR_EVAL: A Transparent Implementation of Common MIR Metrics
复制标题

DOI:
--
复制
发表时间:
2014
期刊:
--
影响因子:
--
通讯作者:
Colin Raffel;Brian McFee;Eric J. Humphrey;J. Salamon;Oriol Nieto;Dawen Liang;D. Ellis
Colin Raffel;Brian McFee;Eric J. Humphrey;J. Salamon;Oriol Nieto;Dawen Liang;D. Ellis
中科院分区:
其他
文献类型:
--
作者:
Colin Raffel;Brian McFee;Eric J. Humphrey;J. Salamon;Oriol Nieto;Dawen Liang;D. Ellis

文献摘要

被引文献

相似文献

MIR 研究领域的核心是评估用于从音乐数据中提取信息的算法。我们推出了 mir_eval,这是一个开源软件库,它提供了用于测量 MIR 算法性能的最常见指标的透明且易于使用的实现。在本文中,我们列举了 mir_eval 实现的指标,并将每个指标与现有实现进行定量比较。当 mir_eval 报告的分数与参考分数有很大差异时,我们会详细说明实现中的差异。我们还简要概述了 mir_eval 的架构、设计和预期用途。 1. 评估 MIR 算法 音乐信息检索 (MIR) 的大部分研究涉及处理原始音乐数据以产生语义信息的系统的开发。这些系统的目标通常被定义为尝试复制人类听众在执行相同任务时的表现 [5]。确定系统有效性的一种自然方法可能是人类研究系统产生的输出并判断其正确性。然而,这只会产生主观评级,并且在评估大量音乐语料库的系统输出时也非常耗时。相反,开发客观指标是为了提供一种明确定义的计算分数的方法,该分数表明每个系统输出的正确性。这些指标通常涉及系统输出与已知正确的参考的启发式比较。随着时间的推移,某些指标已成为每个指标的标准*请直接通信至 craffel@gmail.com c © Colin Raffel、Brian McFee、Eric J. Humphrey、Justin Salamon、Oriol Nieto、Dawen Liang、Daniel P. W. Ellis。根据知识共享署名 4.0 国际许可证 (CC BY 4.0) 获得许可。署名:科林·拉斐尔、布莱恩·麦克菲、埃里克·J·汉弗莱、贾斯汀·萨拉蒙、奥里奥尔·涅托、梁达文、丹尼尔·P·W·埃利斯。
Central to the field of MIR research is the evaluation of algorithms used to extract information from music data. We present mir_eval, an open source software library which provides a transparent and easy-to-use implementation of the most common metrics used to measure the performance of MIR algorithms. In this paper, we enumerate the metrics implemented by mir_eval and quantitatively compare each to existing implementations. When the scores reported by mir_eval differ substantially from the reference, we detail the differences in implementation. We also provide a brief overview of mir_eval’s architecture, design, and intended use. 1. EVALUATING MIR ALGORITHMS Much of the research in Music Information Retrieval (MIR) involves the development of systems that process raw music data to produce semantic information. The goal of these systems is frequently defined as attempting to duplicate the performance of a human listener given the same task [5]. A natural way to determine a system’s effectiveness might be for a human to study the output produced by the system and judge its correctness. However, this would yield only subjective ratings, and would also be extremely timeconsuming when evaluating a system’s output over a large corpus of music. Instead, objective metrics are developed to provide a well-defined way of computing a score which indicates each system’s output’s correctness. These metrics typically involve a heuristically-motivated comparison of the system’s output to a reference which is known to be correct. Over time, certain metrics have become standard for each ∗Please direct correspondence to craffel@gmail.com c © Colin Raffel, Brian McFee, Eric J. Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, Daniel P. W. Ellis. Licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0). Attribution: Colin Raffel, Brian McFee, Eric J. Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, Daniel P. W. Ellis.