Borderline Oversampling in Feature Space for Learning Algorithms in Imbalanced Data Environments

Borderline Oversampling in Feature Space for Learning Algorithms in Imbalanced Data Environments
复制标题

DOI:
--
复制
发表时间:
2016
期刊:
--
影响因子:
--
通讯作者:
Kittipat Savetratanakaree;Kingkarn Sookhanaphibarn;Sarun Intakosum;R. Thawonmas
Kittipat Savetratanakaree;Kingkarn Sookhanaphibarn;Sarun Intakosum;R. Thawonmas
中科院分区:
其他
文献类型:
--
作者:
Kittipat Savetratanakaree;Kingkarn Sookhanaphibarn;Sarun Intakosum;R. Thawonmas

文献摘要

相似文献

为了提高支持向量机在非平衡数据环境中的性能,提出了一种利用特征空间中的欧氏距离对新的少数类实例进行边界过采样的新方法。支持向量机是一种非常成功的分类器,在假设类数据分布均衡的各种应用中都是如此。然而,在处理多数类实例远远多于少数类实例的不平衡数据集时,支持向量机是无效的。我们提出的基于特征空间的边界过采样方法,能够有效地处理不平衡数据,识别新的少数类实例,从而提高支持向量机的分类性能。使用该方法进行的类预测实验结果表明,在g-均值和F-度量方面,该方法比现有的Smote、边界Smote和边界过采样方法具有更好的性能。
In this paper, we propose a new approach to over-sample new minority-class instances along the borderline using the Euclidean distance in the feature space to improve support vector machine (SVM) performance in imbalanced data environments. SVM has been an outstandingly successful classifier in a wide variety of applications where balanced class data distribution is assumed. However, SVM is ineffective when coping with imbalanced datasets whereby the majorityclass instances far outnumber the minority-class instances. Our new approach, called Borderline Over-sampling in the Feature Space, can deal with imbalanced data to effectively recognize new minority-class instances for better classification with SVM. The results of our class prediction experiments using the proposed approach demonstrate better performance than the existing SMOTE, Borderline-SMOTE and borderline over-sampling methods in terms of the g-mean and F-measure.