Margin calibration in SVM class-imbalanced learning

Margin calibration in SVM class-imbalanced learning
复制标题

DOI:
10.1016/j.neucom.2009.08.006
复制
发表时间:
2009-12
期刊:
影响因子:
6
通讯作者:
Chan-Yun Yang;Jr-Syu Yang;Jianjun Wang
Chan-Yun Yang;Jr-Syu Yang;Jianjun Wang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Chan-Yun Yang;Jr-Syu Yang;Jianjun Wang

文献摘要

被引文献

相似文献

不平衡数据集学习是机器学习中的一个重要实际问题,甚至在支持向量机(SVM)中也是如此。在这项研究中,一个众所周知的参考模型,用于解决由Veropoulos等人提出的问题,首先是研究。从损失函数的角度出发,将参考代价敏感原型识别为惩罚正则化模型。直观地说,损失函数不仅可以改变惩罚,而且可以改变边缘,以恢复有偏的决策边界。本研究主要集中在从边缘的影响,然后扩展到一个更一般的修改模型。如原型中所提出的,修改首先采用逆比例正则化惩罚来重新加权不平衡类。除了惩罚正则化之外,该修改还采用了一种边缘补偿来导致边缘不平衡,这使得决策边界漂移。两个正则化因子,惩罚和利润,因此,建议实现无偏分类。与惩罚正则化相关联的边缘补偿在这里被用来校准和细化有偏的决策边界,以进一步减小偏差。使用受试者工作特征曲线下面积(AuROC)检查性能,即使参考模型实现了最佳性能,修改也显示出比参考模型相对更高的分数。还包括一些有用的特性发现经验,这可能是方便的未来的应用。所有的理论描述和实验验证表明,该模型的潜力,竞争高度无偏的准确性,在一个复杂的不平衡数据集。
Imbalanced dataset learning is an important practical issue in machine learning, even in support vector machines (SVMs). In this study, a well known reference model for solving the problem proposed by Veropoulos et al., is first studied. From the aspect of loss function, the reference cost sensitive prototype is identified as a penalty-regularized model. Intuitively, the loss function can change not only the penalty but also the margin to recover the biased decision boundary. This study focuses mainly on the effect from the margin and then extends the model to a more general modification. As proposed in the prototype, the modification first adopts an inversed proportional regularized penalty to re-weight the imbalanced classes. In addition to the penalty regularization, the modification then employs a margin compensation to lead the margin to be lopsided, which enables the decision boundary drift. Two regularization factors, the penalty and margin, are hence suggested for achieving an unbiased classification. The margin compensation, associating with the penalty regularization, is here utilized to calibrate and refine the biased decision boundary to further reduce the bias. With the area under the receiver operating characteristic curve (AuROC) for examining the performance, the modification shows relative higher scores than the reference model, even though the optimal performance is achieved by the reference model. Some useful characteristics found empirically are also included, which may be convenient for the future applications. All the theoretical descriptions and experimental validations show the proposed model's potential to compete for highly unbiased accuracy in a complex imbalanced dataset.