Improving identification of difficult small classes by balancing class distribution

Improving identification of difficult small classes by balancing class distribution
复制标题

DOI:
10.1007/3-540-48229-6_9
复制
发表时间:
2001-01-01
期刊:
ARTIFICIAL INTELLIGENCE IN MEDICINE, PROCEEDINGS
影响因子:
--
通讯作者:
Laurikkala, J
Laurikkala, J
中科院分区:
其他
文献类型:
--
作者:
Laurikkala, J

文献摘要

被引文献

相似文献

我们研究了三种方法,通过平衡不平衡的类分布和数据约简来提高对困难小类的识别。在10个数据集的实验中,新的邻域清洗规则(NCL)方法的性能优于简单的随机和片面选择方法。所有的缩减方法都提高了对小类的识别率(20%-30%),但差异不显著。然而,在准确性、真阳性率和真阴性率方面,用3最近邻法和C4.5方法获得的准确性和真阴性率的显著差异有利于NCL。结果表明,NCL是一种有用的方法,可以改进困难小类的建模,并构建分类器来从真实世界的数据中识别这些类。
We studied three methods to improve identification of difficult small classes by balancing imbalanced class distribution with data reduction. The new method, neighborhood cleaning rule (NCL), outperformed simple random and one-sided selection methods in experiments with ten data sets. All reduction methods improved identification of small classes (20-30%), but the differences were insignificant. However, significant differences in accuracies, true-positive rates and true-negative rates obtained with the 3-nearest neighbor method and C4.5 from the reduced data favored NCL. The results suggest that NCL is a useful method for improving the modeling of difficult small classes, and for building classifiers to identify these classes from the real-world data.