A combination of clustering-based under-sampling with ensemble methods for solving imbalanced class problem in intelligent systems

A combination of clustering-based under-sampling with ensemble methods for solving imbalanced class problem in intelligent systems
复制标题

DOI:
10.1016/j.techfore.2021.120796
复制
发表时间:
2021-05-11
影响因子:
12
通讯作者:
Palmieri, Francesco
Palmieri, Francesco
中科院分区:
管理学1区
文献类型:
--
作者:
Shahabadi, Mohammad Saleh Ebrahimi;Tabrizchi, Hamed;Palmieri, Francesco

文献摘要

被引文献

相似文献

如今,大多数真实世界的数据集都存在数据样本在类中分布不平衡的问题,特别是当代表较大类(大多数)的数据数量远远大于较小类(少数)时。为了解决这个问题,已经提出了各种类型的欠采样或过采样技术,以通过分别减少或增加多数类或少数类中的样本数量来创建每个类中具有相等数量的样本的数据集。包围式分类器使用多种学习算法来提高分类的准确性。在此基础上,将欠采样或过采样方法与集成分类器相结合,可以得到性能更好的模型。通过使用聚类和新的欠采样方法,本研究旨在提出一种新的基于聚类的欠采样方法来创建一个平衡的数据集。该方法采用k-means聚类算法对数据进行聚类,用马氏距离分析每个类中样本到质心的距离,并采用保留每个类中数据分布模式的选择方法。对于从KEEL库中的44个基准数据集获得的实验结果,所提出的方法表现优于七个国家的最先进的方法。
Nowadays, most real-world datasets suffer from the problem of imbalanced distribution of data samples in classes, especially when the number of data representing the larger class (majority) is much greater than that of the smaller class (minority). In order to solve this problem, various types of undersampling or oversampling techniques have been proposed to create a dataset with equal number of samples in each class by reducing or increasing the number of samples in majority or minority classes, respectively. Ensemble classifiers use multiple learning algorithms to enhance the accuracy of classification. Based on the results, combining undersampling or oversampling methods with ensemble classifiers can result in models with better performance. By using both clustering and new undersampling methods, the present study aimed to propose a novel clustering-based undersampling method to create a balanced dataset. This method uses k-means clustering algorithm for clustering the data, Mahalanobis distance to analyze samples distance in each cluster to centroid, and a selection method that preserves the pattern of data distribution in each cluster. Regarding the experimental results obtained by 44 benchmark datasets from KEEL repository, the proposed approach performed better than that of seven state-of-the-art approaches.