Precision-Recall versus Accuracy and the Role of Large Data Sets

Precision-Recall versus Accuracy and the Role of Large Data Sets
复制标题

DOI:
10.1609/aaai.v33i01.33014039
复制
发表时间:
2019-07
期刊:
--
影响因子:
--
通讯作者:
Brendan Juba;Hai S. Le
Brendan Juba;Hai S. Le
中科院分区:
其他
文献类型:
--
作者:
Brendan Juba;Hai S. Le

文献摘要

被引文献

相似文献

长期以来,数据挖掘和机器学习的实践者已经观察到,数据集中的类不平衡会对对该数据进行培训的分类器的质量产生负面影响。已经提出了许多应对这种不平衡的技术,但几乎所有的理论基础几乎都缺乏。相比之下,机器学习的标准理论分析完全不依赖类的失衡。统计学习的基本定理确定了估计分类器的准确性所需的示例数量,其复杂性(VC维度)的函数以及所需的信心;班级不平衡不会在任何地方输入这些公式。在这项工作中,我们从精确和召回方面考虑了分类器性能的度量,该措施被广泛建议更适合于对不平衡数据的分类。我们观察到,每当精确度适中大时,精度和回忆的较差就在阶级不平衡加权准确性的小恒定因子之内。这种观察结果的必然是,需要大量示例需要解决阶级失衡,这一发现我们也从经验上说明了这一发现。
Practitioners of data mining and machine learning have long observed that the imbalance of classes in a data set negatively impacts the quality of classifiers trained on that data. Numerous techniques for coping with such imbalances have been proposed, but nearly all lack any theoretical grounding. By contrast, the standard theoretical analysis of machine learning admits no dependence on the imbalance of classes at all. The basic theorems of statistical learning establish the number of examples needed to estimate the accuracy of a classifier as a function of its complexity (VC-dimension) and the confidence desired; the class imbalance does not enter these formulas anywhere. In this work, we consider the measures of classifier performance in terms of precision and recall, a measure that is widely suggested as more appropriate to the classification of imbalanced data. We observe that whenever the precision is moderately large, the worse of the precision and recall is within a small constant factor of the accuracy weighted by the class imbalance. A corollary of this observation is that a larger number of examples is necessary and sufficient to address class imbalance, a finding we also illustrate empirically.