Generating ground truth for music mood classification using mechanical turk

Generating ground truth for music mood classification using mechanical turk
复制标题

使用 Mechanical Turk 生成音乐情绪分类的基本事实

DOI:
10.1145/2232817.2232842
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
Xiao Hu
Xiao Hu
中科院分区:
--
文献类型:
--
作者:
Jin Ha Lee;Xiao Hu

文献摘要

参考文献

被引文献

相似文献

情绪是音乐数字图书馆和在线音乐库的一个重要接入点,但为评估各种音乐情绪分类算法生成基础事实是一个具有挑战性的问题。这是因为,由于音乐情绪的主观性,收集足够的人类判断是费时且昂贵的。在本研究中,我们探索了使用Amazon Mechanical Turk (MTurk)众包音乐情绪分类判断的可行性。具体来说,我们比较了为年度音乐信息检索评价交换(MIREX)收集的情绪分类判断与使用MTurk收集的判断。我们的数据显示,MIREX和MTurk的情绪集群的总体分布和一致性率具有可比性。然而,土耳其人倾向于比MIREX评估者更不同意预先标记的情绪集群。使用两组数据生成的系统评估结果除了使用Friedman检验检测到一个统计显著对外,大部分是相同的。我们得出的结论是,MTurk可以作为地面真相收集的可行替代方案,但对于特定的情绪集群有一些保留。
Mood is an important access point in music digital libraries and online music repositories, but generating ground truth for evaluating various music mood classification algorithms is a challenging problem. This is because collecting enough human judgments is time-consuming and costly due to the subjectivity of music mood. In this study, we explore the viability of crowdsourcing music mood classification judgments using Amazon Mechanical Turk (MTurk). Specifically, we compare the mood classification judgments collected for the annual Music Information Retrieval Evaluation eXchange (MIREX) with judgments collected using MTurk. Our data show that the overall distribution of mood clusters and agreement rates from MIREX and MTurk were comparable. However, Turkers tended to agree less with the pre-labeled mood clusters than MIREX evaluators. The system evaluation results generated using both sets of data were mostly the same except for detecting one statistically significant pair using Friedman's test. We conclude that MTurk can potentially serve as a viable alternative for ground truth collection, with some reservation with regards to particular mood clusters.
使用 Zig-zag 产品进行图扩展时几乎处处一致
DOI: --
发表时间: 2017
期刊:
影响因子: --
作者:
Hans L. Bodlaender;Hirotaka Ono;Yota Otachi;切上太希,小柴健史
通讯作者: 切上太希,小柴健史