Cluster-based under-sampling approaches for imbalanced data distributions

Cluster-based under-sampling approaches for imbalanced data distributions
复制标题

DOI:
10.1016/j.eswa.2008.06.108
复制
发表时间:
2009-04
期刊:
Expert Syst. Appl.
影响因子:
--
通讯作者:
Show-Jane Yen;Yue-Shi Lee
Show-Jane Yen;Yue-Shi Lee
中科院分区:
其他
文献类型:
--
作者:
Show-Jane Yen;Yue-Shi Lee

文献摘要

被引文献

相似文献

在分类问题中,训练数据对分类精度有很大的影响。然而,在实际应用中,数据往往是不均衡的类分布,即大多数数据属于多数类,少数数据属于少数类。在这种情况下,如果所有数据都被用作训练数据,则分类器倾向于预测大多数传入数据属于多数类。因此,在不平衡类分布问题中,选择合适的训练数据进行分类是非常重要的。本文提出了基于聚类的欠采样方法来选择具有代表性的数据作为训练数据,以提高少数类的分类精度,并研究了欠采样方法在不平衡类分布环境中的效果。实验结果表明,我们的基于聚类的欠采样方法优于其他欠采样技术在以前的研究。
For classification problem, the training data will significantly influence the classification accuracy. However, the data in real-world applications often are imbalanced class distribution, that is, most of the data are in majority class and little data are in minority class. In this case, if all the data are used to be the training data, the classifier tends to predict that most of the incoming data belongs to the majority class. Hence, it is important to select the suitable training data for classification in the imbalanced class distribution problem. In this paper, we propose cluster-based under-sampling approaches for selecting the representative data as training data to improve the classification accuracy for minority class and investigate the effect of under-sampling methods in the imbalanced class distribution environment. The experimental results show that our cluster-based under-sampling approaches outperform the other under-sampling techniques in the previous studies.