A Novel Boundary Oversampling Algorithm Based on Neighborhood Rough Set Model: NRSBoundary-SMOTE

A Novel Boundary Oversampling Algorithm Based on Neighborhood Rough Set Model: NRSBoundary-SMOTE
复制标题

一种基于邻域粗糙集模型的边界过采样新算法:NRSBoundary-SMOTE

DOI:
10.1155/2013/694809
复制
发表时间:
2013-01-01
影响因子:
--
通讯作者:
Li, Hang
Li, Hang
中科院分区:
工程技术4区
文献类型:
--
作者:
Hu, Feng;Li, Hang

文献摘要

被引文献

相似文献

粗糙集理论是 Pawlak 提出的一种强大的数学工具,用于处理不精确、不确定和模糊的信息。基于邻域的粗糙集模型扩展了粗糙集理论;它可以将数据集分为三个部分。边界区域表示多数类样本和少数类样本重叠。根据我们对原始数据集分布的了解,我们仅对边界区域中与多数类样本重叠的少数类样本进行过采样。因此,NRSBoundary-SMOTE可以扩大少数类别的决策空间;同时,它也会缩小多数阶层的决策空间。在四种分类器上进行实验后,NRSBoundary-SMOTE在使用C4.5、CART和KNN时比其他方法具有更高的精度,但在分类器SVM上比SMOTE差。
Rough set theory is a powerful mathematical tool introduced by Pawlak to deal with imprecise, uncertain, and vague information. The Neighborhood-Based Rough Set Model expands the rough set theory; it could divide the dataset into three parts. And the boundary region indicates that the majority class samples and the minority class samples are overlapped. On the basis of what we know about the distribution of original dataset, we only oversample the minority class samples, which are overlapped with the majority class samples, in the boundary region. So, the NRSBoundary-SMOTE can expand the decision space for the minority class; meanwhile, it will shrink the decision space for the majority class. After conducting an experiment on four kinds of classifiers, NRSBoundary-SMOTE has higher accuracy than other methods when C4.5, CART, and KNN are used but it is worse than SMOTE on classifier SVM.