The use of the area under the roc curve in the evaluation of machine learning algorithms

The use of the area under the roc curve in the evaluation of machine learning algorithms
复制标题

DOI:
10.1016/s0031-3203(96)00142-2
复制
发表时间:
1997-07-01
影响因子:
8
通讯作者:
Bradley, AP
Bradley, AP
中科院分区:
计算机科学1区
文献类型:
--
作者:
Bradley, AP

文献摘要

被引文献

相似文献

在本文中,我们研究了使用受试者工作特征曲线(ROC)下的面积(AUC)作为机器学习算法的性能指标。作为一个案例研究,我们在六个“真实世界”的医疗诊断数据集上评估了六种机器学习算法(C4.5,多尺度分类器,感知器,多层感知器,k最近邻和二次判别函数)。我们比较和讨论使用AUC的更传统的整体准确性,并发现AUC表现出一些理想的性能相比,整体准确性:增加的灵敏度在方差分析(ANOVA)测试;标准误,降低AUC和测试样本的数量增加;决策阈值独立;它是不变的先验类概率。本文最后建议优先使用AUC而不是整体准确度来评估机器学习算法的“单一数字”。(C)1997年模式识别学会。
In this paper we investigate the use of the area under the receiver operating characteristic (ROC) curve (AUC) as a performance measure for machine learning algorithms. As a case study we evaluate six machine learning algorithms (C4.5, Multiscale Classifier, Perceptron, Multi-layer Perceptron, k-Nearest Neighbours, and a Quadratic Discriminant Function) on six ''real world'' medical diagnostics data sets. We compare and discuss the use of AUC to the more conventional overall accuracy and find that AUC exhibits a number of desirable properties when compared to overall accuracy: increased sensitivity in Analysis of Variance (ANOVA) tests; a standard error that decreased as both AUC and the number of test samples increased; decision threshold independent; and it is invariant to a priori class probabilities. The paper concludes with the recommendation that AUC be used in preference to overall accuracy for ''single number'' evaluation of machine learning algorithms. (C) 1997 Pattern Recognition Society.