The Matthews Correlation Coefficient (MCC) is More Informative Than Cohen's Kappa and Brier Score in Binary Classification Assessment

The Matthews Correlation Coefficient (MCC) is More Informative Than Cohen's Kappa and Brier Score in Binary Classification Assessment
复制标题

DOI:
10.1109/access.2021.3084050
复制
发表时间:
2021-01-01
期刊:
影响因子:
3.9
通讯作者:
Jurman, Giuseppe
Jurman, Giuseppe
中科院分区:
计算机科学3区
文献类型:
--
作者:
Chicco, Davide;Warrens, Matthijs J.;Jurman, Giuseppe

文献摘要

被引文献

相似文献

即使测量二进制分类的结果是机器学习和统计学中的一项关键任务,但对于为此使用哪种统计率尚未达成共识。在上个世纪,计算机科学和统计学社区已经引入了几个分数来总结关于地面真值的预测的正确性。在这些评分中,马修斯相关系数(MCC)被证明比混淆熵、准确性、F-1评分、平衡准确性、博彩公司信息、显著性和诊断比值比具有几个优势:事实上,只有当大多数预测的负数据实例和大多数正数据实例是正确的时,MCC才产生高分,因此,它在不平衡的数据集上非常值得信赖。在这项研究中,我们比较MCC与其他两个流行的分数:科恩的Kappa,一个度量,起源于社会科学,和Brier分数,一个严格正确的评分功能出现在天气预报研究。在解释了MCC与这两个比率之间的数学属性和关系之后,我们报告了一些用例,其中这些分数生成不同的值,这导致不一致的结果,其中MCC提供了更真实和信息量更大的结果。我们强调的原因,它是更可取的使用MCC,而不是科恩的Kappa和Brier评分来评估二进制分类。
Even if measuring the outcome of binary classifications is a pivotal task in machine learning and statistics, no consensus has been reached yet about which statistical rate to employ to this end. In the last century, the computer science and statistics communities have introduced several scores summing up the correctness of the predictions with respect to the ground truth values. Among these scores, the Matthews correlation coefficient (MCC) was shown to have several advantages over confusion entropy, accuracy, F-1 score, balanced accuracy, bookmaker informedness, markedness, and diagnostic odds ratio: MCC, in fact, produces a high score only if the majority of the predicted negative data instances and the majority of the positive data instances are correct, and therefore it results being very trustworthy on imbalanced datasets. In this study, we compare MCC with two other popular scores: Cohen's Kappa, a metric that originated in social sciences, and the Brier score, a strictly proper scoring function which emerged in weather forecasting studies. After explaining the mathematical properties and the relationships between MCC and each of these two rates, we report some use cases where these scores generate different values, which lead to discordant outcomes, where MCC provides a more truthful and informative result. We highlight the reasons why it is more advisable to use MCC rather that Cohen's Kappa and the Brier score to evaluate binary classifications.