Data Mining for Imbalanced Datasets: An Overview

Data Mining for Imbalanced Datasets: An Overview
复制标题

DOI:
10.1007/978-0-387-09823-4_45
复制
发表时间:
2010-01-01
期刊:
DATA MINING AND KNOWLEDGE DISCOVERY HANDBOOK, SECOND EDITION
影响因子:
--
通讯作者:
Chawla, Nitesh V.
Chawla, Nitesh V.
中科院分区:
其他
文献类型:
--
作者:
Chawla, Nitesh V.

文献摘要

被引文献

相似文献

如果分类类别没有近似相等地表示,则数据集是不平衡的。近年来,人们越来越关注将机器学习技术应用于困难的“现实世界”问题,其中许多问题的特征是不平衡的数据。此外,测试数据的分布可能与训练数据的分布不同,并且真实的误分类成本在学习时可能是未知的。预测准确性是评估分类器性能的一个流行选择,但当数据不平衡和/或不同错误的成本显著变化时,预测准确性可能不合适。在本章中,我们将讨论一些用于平衡数据集的抽样技术,以及更适合挖掘不平衡数据集的性能度量。
A dataset is imbalanced if the classification categories are not approximately equally represented. Recent years brought increased interest in applying machine learning techniques to difficult "real-world" problems, many of which are characterized by imbalanced data. Additionally the distribution of the testing data may differ from that of the training data, and the true misclassification costs may be unknown at learning time. Predictive accuracy, a popular choice for evaluating performance of a classifier, might not be appropriate when the data is imbalanced and/or the costs of different errors vary markedly. In this Chapter, we discuss some of the sampling techniques used for balancing the datasets, and the performance measures more appropriate for mining imbalanced datasets.