A multiple resampling method for learning from imbalanced data sets

A multiple resampling method for learning from imbalanced data sets
复制标题

DOI:
10.1111/j.0824-7935.2004.t01-1-00228.x
复制
发表时间:
2004-02-01
影响因子:
2.8
通讯作者:
Japkowicz, N
Japkowicz, N
中科院分区:
计算机科学4区
文献类型:
--
作者:
Estabrooks, A;Jo, TH;Japkowicz, N

文献摘要

被引文献

相似文献

重采样方法是处理类不平衡问题的常用方法。与其他方法相比,它们的优势在于它们是外部的,因此很容易运输。尽管这样的方法可以非常简单地实现,但最有效地调优它们并不是一项容易的任务。特别是,不清楚过采样是否比欠采样更有效,以及应该使用哪种过采样或欠采样比率。本文对这些问题进行了实验研究,并得出结论:结合重采样方法的不同表达式是解决调整问题的有效方法。在路透社-21578文本集合的不平衡子集上对所提出的组合方案进行了评估,结果表明该组合方案对于这些问题是非常有效的。
Resampling methods are commonly used for dealing with the class-imbalance problem. Their advantage over other methods is that they are external and thus, easily transportable. Although such approaches can be very simple to implement, tuning them most effectively is not an easy task. In particular, it is unclear whether oversampling is more effective than undersampling and which oversampling or undersampling rate should be used. This paper presents an experimental study of these questions and concludes that combining different expressions of the resampling approach is an effective solution to the tuning problem. The proposed combination scheme is evaluated on imbalanced subsets of the Reuters-21578 text collection and is shown to be quite effective for these problems.