Applying support vector machines to imbalanced datasets

Applying support vector machines to imbalanced datasets
复制标题

DOI:
10.1007/978-3-540-30115-8_7
复制
发表时间:
2004-01-01
期刊:
MACHINE LEARNING: ECML 2004, PROCEEDINGS
影响因子:
--
通讯作者:
Japkowicz, N
Japkowicz, N
中科院分区:
其他
文献类型:
--
作者:
Akbani, R;Kwek, S;Japkowicz, N

文献摘要

被引文献

相似文献

支持向量机 (SVM) 已得到广泛研究,并在许多应用中取得了显着的成功。然而,当 SVM 应用于从不平衡数据集中学习的问题时,它的成功非常有限,其中负面实例的数量远远超过正面实例(例如,在基因分析和检测信用卡欺诈中)。本文讨论了这种失败背后的因素,并解释了为什么对训练数据进行欠采样的常见策略可能不是 SVM 的最佳选择。然后,我们提出了一种克服这些问题的算法,该算法基于 Chawla 等人的 SMOTE 算法的变体,并结合 Veropoulos 等人的不同错误成本算法。我们将我们的算法与这两种算法以及欠采样和常规 SVM 的性能进行比较,并表明我们的算法优于所有算法。
Support Vector Machines (SVM) have been extensively studied and have shown remarkable success in many applications. However the success of SVM is very limited when it is applied to the problem of learning from imbalanced datasets in which negative instances heavily outnumber the positive instances (e.g. in gene profiling and detecting credit card fraud). This paper discusses the factors behind this failure and explains why the common strategy of undersampling the training data may not be the best choice for SVM. We then propose an algorithm for overcoming these problems which is based on a variant of the SMOTE algorithm by Chawla et al, combined with Veropoulos et al's different error costs algorithm. We compare the performance of our algorithm against these two algorithms, along with undersampling and regular SVM and show that our algorithm outperforms all of them.