Classifying imbalanced data in distance-based feature space

Classifying imbalanced data in distance-based feature space
复制标题

DOI:
10.1007/s10115-015-0846-3
复制
发表时间:
2016-03
影响因子:
2.7
通讯作者:
S. Ando
S. Ando
中科院分区:
计算机科学4区
文献类型:
--
作者:
S. Ando

文献摘要

被引文献

相似文献

类不平衡是实际分类问题中的一个重要问题。重要的对策,如重新采样,实例加权,成本敏感的学习已经开发出来,但也有各自的方法的局限性和优势。合成重采样方法具有广泛的适用性,但需要矢量表示来生成额外的实例。基于实例的方法可以应用于距离空间数据,但对于全局目标不容易处理。代价敏感学习可以最小化给定错误代价的期望代价,但通常不扩展到非线性度量,如F-度量和曲线下面积。为了解决上述问题,本文提出了一种最近邻分类模型,该模型采用类加权方案来抵消类不平衡,并采用凸优化技术来学习其权重参数。因此,所提出的模型保持了简单的基于实例的预测规则,但保留了学习的数学支持,以最大限度地提高训练集的非线性性能指标。通过实验研究,评估了该算法在不平衡距离空间数据上的性能,并与现有方法进行了比较。
Class imbalance is a significant issue in practical classification problems. Important countermeasures, such as re-sampling, instance-weighting, and cost-sensitive learning have been developed, but there are limitations as well as advantages to respective approaches. The synthetic re-sampling methods have wide applicability, but require a vector representation to generate additional instances. The instance-based methods can be applied to distance space data, but are not tractable with regard to a global objective. The cost-sensitive learning can minimize the expected cost given the costs of error, but generally does not extend to nonlinear measures, such as F-measure and area under the curve. In order to address the above shortcomings, this paper proposes a nearest neighbor classification model which employs a class-wise weighting scheme to counteract the class imbalance and a convex optimization technique to learn its weight parameters. As a result, the proposed model maintains the simple instance-based rule for prediction, yet retains a mathematical support for learning to maximize a nonlinear performance measure over the training set. An empirical study is conducted to evaluate the performance of the proposed algorithm on the imbalanced distance space data and make comparison with existing methods.