SVMs Modeling for Highly Imbalanced Classification

SVMs Modeling for Highly Imbalanced Classification
复制标题

DOI:
10.1109/tsmcb.2008.2002909
复制
发表时间:
2009-02-01
影响因子:
--
通讯作者:
Krasser, Sven
Krasser, Sven
中科院分区:
其他
文献类型:
--
作者:
Tang, Yuchun;Zhang, Yan-Qing;Krasser, Sven

文献摘要

被引文献

相似文献

传统的分类算法在高度不平衡的数据集上的性能有限。一个流行的工作流,以对付阶级不平衡的问题一直是应用程序的抽样策略。在这封信中,我们专注于设计修改支持向量机(SVM),以适当地解决类不平衡的问题。我们在SVM建模中采用了不同的“再平衡”算法,包括成本敏感学习,以及过采样和欠采样。这些基于SVM的策略进行了比较与各种国家的最先进的方法在各种数据集上使用各种指标,包括G-均值,面积下的接收器操作特征曲线,F-措施,和面积下的精度/召回曲线。我们表明,我们能够超越或匹配以前已知的最佳算法对每个数据集。特别是,在这种对应关系中考虑的四种SVM变体中,新的粒度SVM重复欠采样算法(GSVM-RU)在有效性和效率方面都是最好的。GSVM-RU是有效的,因为它可以最大限度地减少信息丢失的负面影响,同时最大限度地提高欠采样过程中数据清洗的积极影响。GSVM-RU通过提取更少的支持向量而有效,因此大大加快了SVM预测。
Traditional classification algorithms can be limited in their performance on highly unbalanced data sets. A popular stream of work for countering the problem of class imbalance has been the application of a sundry of sampling strategies. In this correspondence, we focus on designing modifications to support vector machines (SVMs) to appropriately tackle the problem of class imbalance. We incorporate different "rebalance" heuristics in SVM modeling, including cost-sensitive learning, and over- and undersampling. These SVM-based strategies are compared with various state-of-the-art approaches on a variety of data sets by using various metrics, including G-mean, area under the receiver operating characteristic curve, F-measure, and area under the precision/recall curve. We show that we am able to surpass or match the previously known best algorithms on each data set. In particular, of the four SVM variations considered in this correspondence, the novel granular SVMs-repetitive undersampling algorithm (GSVM-RU) is the best in terms of both effectiveness and efficiency. GSVM-RU is effective, as it can minimize the negative effect of information loss while maximizing the positive effect of data cleaning in the undersampling process. GSVM-RU is efficient by extracting much less support vectors and, hence, greatly speeding up SVM prediction.