The effect of class distribution on classifier learning: an empirical study

The effect of class distribution on classifier learning: an empirical study
复制标题

DOI:
10.7282/t3-vpfw-sf95
复制
发表时间:
2001-08
期刊:
--
影响因子:
--
通讯作者:
Gary M. Weiss;F. Provost
Gary M. Weiss;F. Provost
中科院分区:
其他
文献类型:
--
作者:
Gary M. Weiss;F. Provost

文献摘要

被引文献

相似文献

在这篇文章中,我们分析了类分布对分类器学习的影响。我们开始描述类分布影响学习的不同方式,以及它如何影响学习分类器的评估。然后,我们提出了两个全面的实验研究的结果。第一项研究比较了从不平衡数据集生成的分类器的性能与从相同数据集的平衡版本生成的分类器的性能。这种比较使我们能够隔离和量化训练集的类分布对学习的影响,并对比分类器在少数类和多数类上的性能。第二项研究评估了什么样的分布是“最好的”训练,关于两个性能指标:分类准确性和ROC曲线下面积(AUC)。分类器归纳的许多研究背后的一个默认假设是,训练数据的类分布应该与数据的“自然”分布相匹配。这项研究表明,自然发生的类分布往往不是最好的学习,往往可以通过使用不同的类分布获得更好的性能。了解分类器性能如何受到类分布的影响可以帮助从业者选择训练数据-在现实世界中,由于计算成本或与采购和准备数据相关的成本,训练示例的数量通常必须受到限制。
In this article we analyze the effect of class distribution on classifier learning. We begin by describing the different ways in which class distribution affects learning and how it affects the evaluation of learned classifiers. We then present the results of two comprehensive experimental studies. The first study compares the performance of classifiers generated from unbalanced data sets with the performance of classifiers generated from balanced versions of the same data sets. This comparison allows us to isolate and quantify the effect that the training set’s class distribution has on learning and contrast the performance of the classifiers on the minority and majority classes. The second study assesses what distribution is "best" for training, with respect to two performance measures: classification accuracy and the area under the ROC curve (AUC). A tacit assumption behind much research on classifier induction is that the class distribution of the training data should match the “natural” distribution of the data. This study shows that the naturally occurring class distribution often is not best for learning, and often substantially better performance can be obtained by using a different class distribution. Understanding how classifier performance is affected by class distribution can help practitioners to choose training data—in real-world situations the number of training examples often must be limited due to computational costs or the costs associated with procuring and preparing the data.