A dynamic over-sampling procedure based on sensitivity for multi-class problems

A dynamic over-sampling procedure based on sensitivity for multi-class problems
复制标题

DOI:
10.1016/j.patcog.2011.02.019
复制
发表时间:
2011-08-01
影响因子:
8
通讯作者:
Antonio Gutierrez, Pedro
Antonio Gutierrez, Pedro
中科院分区:
计算机科学1区
文献类型:
--
作者:
Fernandez-Navarro, Francisco;Hervas-Martinez, Cesar;Antonio Gutierrez, Pedro

文献摘要

被引文献

相似文献

不平衡数据集分类对机器学习框架下的研究提出了新的挑战。当表示数据集的一个类(通常是感兴趣的概念)的模式数量远低于其他类时,就会出现这个问题。因此,学习模型必须适应这种情况,这在实际应用程序中非常常见。本文提出了一种动态过采样方法,用于改进两类以上不平衡数据集的分类。这一过程被纳入模因算法(MA),优化径向基函数神经网络(RBFNNs)。为了解决类不平衡问题,训练数据分两个阶段进行重采样。在第一阶段,对少数类应用过采样程序,以部分平衡类的大小。然后,运行MA并在进化的不同代中对数据进行过采样,生成最小灵敏度类(总体中最佳RBFNN精度最差的类)的新模式。使用13个不平衡基准分类数据集对所提出的方法进行了测试,这些数据集来自知名的机器学习问题和一个复杂的微生物生长问题。并将其与其他专门用于处理不平衡数据的神经网络方法进行了比较。这些方法包括预处理阶段的不同过采样过程,一种阈值移动方法,其中输出阈值移动到便宜的类和集成方法,结合使用这些技术获得的模型。结果表明,本文提出的方法能够提高泛化集的灵敏度,对每个类别都能获得较高的准确率水平和较好的分类水平。(C) 2011 Elsevier Ltd.版权所有。
Classification with imbalanced datasets supposes a new challenge for researches in the framework of machine learning. This problem appears when the number of patterns that represents one of the classes of the dataset (usually the concept of interest) is much lower than in the remaining classes. Thus, the learning model must be adapted to this situation, which is very common in real applications. In this paper, a dynamic over-sampling procedure is proposed for improving the classification of imbalanced datasets with more than two classes. This procedure is incorporated into a memetic algorithm (MA) that optimizes radial basis functions neural networks (RBFNNs). To handle class imbalance, the training data are resampled in two stages. In the first stage, an over-sampling procedure is applied to the minority class to balance in part the size of the classes. Then, the MA is run and the data are over-sampled in different generations of the evolution, generating new patterns of the minimum sensitivity class (the class with the worst accuracy for the best RBFNN of the population). The methodology proposed is tested using 13 imbalanced benchmark classification datasets from well-known machine learning problems and one complex problem of microbial growth. It is compared to other neural network methods specifically designed for handling imbalanced data. These methods include different over-sampling procedures in the preprocessing stage, a threshold-moving method where the output threshold is moved toward inexpensive classes and ensembles approaches combining the models obtained with these techniques. The results show that our proposal is able to improve the sensitivity in the generalization set and obtains both a high accuracy level and a good classification level for each class. (C) 2011 Elsevier Ltd. All rights reserved.