Data Mining for Imbalanced Datasets: An Overview
Data Mining for Imbalanced Datasets: An Overview
复制标题
DOI:
10.1007/978-0-387-09823-4_45
复制
发表时间:
2010-01-01
期刊:
影响因子:
--
通讯作者:
Chawla, Nitesh V.
中科院分区:
文献类型:
--
作者:
Chawla, Nitesh V.
A dataset is imbalanced if the classification categories are not approximately equally represented. Recent years brought increased interest in applying machine learning techniques to difficult "real-world" problems, many of which are characterized by imbalanced data. Additionally the distribution of the testing data may differ from that of the training data, and the true misclassification costs may be unknown at learning time. Predictive accuracy, a popular choice for evaluating performance of a classifier, might not be appropriate when the data is imbalanced and/or the costs of different errors vary markedly. In this Chapter, we discuss some of the sampling techniques used for balancing the datasets, and the performance measures more appropriate for mining imbalanced datasets.