Correlation analysis of performance measures for multi-label classification

Correlation analysis of performance measures for multi-label classification
复制标题

DOI:
10.1016/j.ipm.2018.01.002
复制
发表时间:
2018-05-01
影响因子:
8.6
通讯作者:
Merschmann, Luiz H. C.
Merschmann, Luiz H. C.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Pereira, Rafael B.;Plastino, Alexandre;Merschmann, Luiz H. C.

文献摘要

被引文献

相似文献

在文本分类、场景分类、生物分子分析和医疗诊断等许多重要应用领域,示例自然与多个类别标签相关联,从而产生多标签分类问题。近年来,这一事实导致了大量的多标签分类研究。为了评估和比较多标签分类器,研究人员已经从单标签范式中调整了评估措施,如精确度和召回率;并且还开发了许多专门用于多标签范式的不同措施,如汉明损失和子集准确度。然而,这些评价措施已被任意使用在多标签分类实验中,没有相关性或偏差的客观分析。这可能会导致误导性的结论,因为实验结果可能会根据所选择的测量子集而有利于特定的行为。此外,由于该领域的不同论文目前采用不同的测量子集,因此很难比较论文之间的结果。在这项工作中,我们提供了一个深入的分析多标签的评价措施,我们给出了具体的建议,研究人员作出明智的决定时,选择多标签分类的评价措施。
In many important application domains, such as text categorization, scene classification, biomolecular analysis and medical diagnosis, examples are naturally associated with more than one class label, giving rise to multi-label classification problems. This fact has led, in recent years, to a substantial amount of research in multi-label classification. In order to evaluate and compare multi-label classifiers, researchers have adapted evaluation measures from the single-label paradigm, like Precision and Recall; and also have developed many different measures specifically for the multi-label paradigm, like Hamming Loss and Subset Accuracy. However, these evaluation measures have been used arbitrarily in multi-label classification experiments, without an objective analysis of correlation or bias. This can lead to misleading conclusions, as the experimental results may appear to favor a specific behavior depending on the subset of measures chosen. Also, as different papers in the area currently employ distinct subsets of measures, it is difficult to compare results across papers. In this work, we provide a thorough analysis of multi label evaluation measures, and we give concrete suggestions for researchers to make an informed decision when choosing evaluation measures for multi-label classification.