RAMOBoost: Ranked Minority Oversampling in Boosting

RAMOBoost: Ranked Minority Oversampling in Boosting
复制标题

DOI:
10.1109/tnn.2010.2066988
复制
发表时间:
2010-10-01
影响因子:
--
通讯作者:
Garcia, Edwardo A.
Garcia, Edwardo A.
中科院分区:
其他
文献类型:
--
作者:
Chen, Sheng;He, Haibo;Garcia, Edwardo A.

文献摘要

被引文献

相似文献

近年来,由于使用和产生不平衡数据的应用程序的爆炸式增长,从不平衡数据中学习引起了学术界和工业界越来越多的关注。然而,由于不平衡数据的复杂特性,许多现实世界的解决方案难以在基于学习的应用程序中提供强大的效率。为了解决这个问题,本文提出了在Boosting(RAMOBoost),这是一个RAMO技术的基础上的集成学习系统中的自适应合成数据生成的想法。简而言之,RAMOBoost根据基于底层数据分布的采样概率分布,在每次学习迭代中自适应地对少数类实例进行排名,并且可以通过使用假设评估过程自适应地将决策边界向难以学习的少数类和多数类实例转移。通过对19个真实数据集的仿真分析,验证了该方法的有效性。
In recent years, learning from imbalanced data has attracted growing attention from both academia and industry due to the explosive growth of applications that use and produce imbalanced data. However, because of the complex characteristics of imbalanced data, many real-world solutions struggle to provide robust efficiency in learning-based applications. In an effort to address this problem, this paper presents Ranked Minority Over-sampling in Boosting (RAMOBoost), which is a RAMO technique based on the idea of adaptive synthetic data generation in an ensemble learning system. Briefly, RAMOBoost adaptively ranks minority class instances at each learning iteration according to a sampling probability distribution that is based on the underlying data distribution, and can adaptively shift the decision boundary toward difficult-to-learn minority and majority class instances by using a hypothesis assessment procedure. Simulation analysis on 19 real-world datasets assessed over various metrics-including overall accuracy, precision, recall, F-measure, G-mean, and receiver operation characteristic analysis-is used to illustrate the effectiveness of this method.