Minority Sub-region Estimation-based Oversampling for Imbalance Learning

Minority Sub-region Estimation-based Oversampling for Imbalance Learning
复制标题

基于少数子区域估计的不平衡学习过采样

DOI:
10.1109/tkde.2020.3010013
复制
发表时间:
2022
影响因子:
8.9
通讯作者:
Wen Zhu
Wen Zhu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yi Sun;Lijun Cai;Bo Liao;Wen Zhu

文献摘要

相似文献

近年来,以偏态分布为特征的班级失衡问题成为一个挑战。许多过采样技术被提出来科普这个问题,其中一些联合收割机将过采样过程与聚类算法相结合,以保证在聚类中产生新的合成样本。然而,由于聚类算法本身的特点,对于距离较远但具有相同少数子区域的样本,通常会被聚类到不同的组中。因此,下面的过采样过程大多是在合成样本不能很好地覆盖完整少数区域的不完整少数子区域中进行的。而现有的算法中,没有一个是直接估计少数子区域的类不平衡问题。因此,一个新的分组算法,命名为方向分布的少数子区域估计(DDMSE),首次提出。该算法利用少数子区域与其他多数子区域相比几乎分布在同一方向的直观观察,巧妙地忽略了聚类算法中距离因子带来的负面影响,估计出少数子区域。最后,在这些少数子区域中生成新的合成样本。在真实数据集上的实验结果表明,该方法的性能与其他过采样方法相当。
Class imbalance problem that characterized with the skew distribution towards the majority arises as one challenge in recent years. Many oversampling techniques have been proposed to cope with this problem and some of them combine the oversampling procedure with the clustering algorithm which guaranteeing new synthetic samples being generated in clusters. However far-away samples but with the same minority sub-region are generally clustered into different groups owing to the characteristic of clustering algorithm itself. Therefore, the following oversampling procedure is mostly carried in incomplete minority sub-regions that synthetic samples not well cover the integral minority region. And to our best knowledge, none of existing algorithm is designed to directly estimate minority sub-regions for class imbalance problem. Thus, one new grouping algorithm, named Direction Distribution-based Minority Sub-region Estimation (DDMSE), is first proposed. The new algorithm exploits the intuitive observation, that the minority with the same sub-region almost distribute within the same direction when compared to other majority, to estimate minority sub-regions that tactfully ignoring negative impacts brought by the distance factor like in clustering algorithms. Finally, new synthetic samples are generated in those minority sub-regions. And experimental results on real-world datasets show the comparable performance with other state-of-the-art oversampling methods.