Measuring classifier performance: a coherent alternative to the area under the ROC curve

Measuring classifier performance: a coherent alternative to the area under the ROC curve
复制标题

DOI:
10.1007/s10994-009-5119-5
复制
发表时间:
2009-10-01
期刊:
影响因子:
7.5
通讯作者:
Hand, David J.
Hand, David J.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hand, David J.

文献摘要

被引文献

相似文献

ROC曲线下面积(AUC)是分类和诊断规则中非常广泛使用的性能度量。它具有客观的吸引力,不需要用户的主观输入。另一方面,AUC也有缺点,其中一些是众所周知的。例如,如果ROC曲线交叉,AUC可能会给出潜在的误导性结果。然而,AUC也有一个更严重的缺陷,一个以前似乎没有被认识到的缺陷。这就是它在错误分类成本方面是根本不连贯的:AUC对不同的分类器使用不同的错误分类成本分布。这意味着使用AUC等同于使用不同的度量来评估不同的分类规则。这相当于说,使用一个分类器,错误分类1类点的严重程度是错误分类0类点的p倍,但是,使用另一个分类器,错误分类1类点的严重程度是p倍,其中p不等于p。这是无意义的,因为不同种类的错误分类的相对严重程度是问题的性质,而不是碰巧选择的分类器。详细探讨了这一特性,并提出了一种简单有效的替代AUC的方法。
The area under the ROC curve (AUC) is a very widely used measure of performance for classification and diagnostic rules. It has the appealing property of being objective, requiring no subjective input from the user. On the other hand, the AUC has disadvantages, some of which are well known. For example, the AUC can give potentially misleading results if ROC curves cross. However, the AUC also has a much more serious deficiency, and one which appears not to have been previously recognised. This is that it is fundamentally incoherent in terms of misclassification costs: the AUC uses different misclassification cost distributions for different classifiers. This means that using the AUC is equivalent to using different metrics to evaluate different classification rules. It is equivalent to saying that, using one classifier, misclassifying a class 1 point is p times as serious as misclassifying a class 0 point, but, using another classifier, misclassifying a class 1 point is P times as serious, where p not equal P. This is nonsensical because the relative severities of different kinds of misclassifications of individual points is a property of the problem, not the classifiers which happen to have been chosen. This property is explored in detail, and a simple valid alternative to the AUC is proposed.