Cost-sensitive boosting for classification of imbalanced data

Cost-sensitive boosting for classification of imbalanced data
复制标题

DOI:
10.1016/j.patcog.2007.04.009
复制
发表时间:
2007-12-01
影响因子:
8
通讯作者:
Wang, Yang
Wang, Yang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Sun, Yamnin;Kamel, Mohamed S.;Wang, Yang

文献摘要

被引文献

相似文献

具有不平衡类分布的数据的分类已经对大多数标准分类器学习算法所能达到的性能造成了显著的缺点,这些标准分类器学习算法假设相对平衡的类分布和相等的误分类成本。班级不平衡问题的显著困难和频繁发生表明需要额外的研究努力。本文的目的是研究适用于大多数分类器学习算法的元技术,旨在提高不平衡数据的分类。AdaBoost算法被报道为一种成功的元技术,用于提高分类精度。通过对AdaBoost算法在解决类不平衡问题方面的优势和不足进行全面分析,探索了三种代价敏感的Boosting算法,并将代价项引入AdaBoost的学习框架。进一步的分析表明,所提出的算法之一符合统计学中的阶段加性模型,以最小化成本指数损失。这些提升算法还研究了他们的加权策略对不同类型的样本,并通过实验在几个真实的世界的医疗数据集,其中类不平衡问题盛行的情况下,他们在识别罕见的情况下的有效性。(c)2007模式识别学会。由爱思唯尔有限公司出版。保留所有权利。
Classification of data with imbalanced class distribution has posed a significant drawback of the performance attainable by most standard classifier learning algorithms, which assume a relatively balanced class distribution and equal misclassification costs. The significant difficulty and frequent occurrence of the class imbalance problem indicate the need for extra research efforts. The objective of this paper is to investigate meta-techniques applicable to most classifier learning algorithms, with the aim to advance the classification of imbalanced data. The AdaBoost algorithm is reported as a successful meta-technique for improving classification accuracy. The insight gained from a comprehensive analysis of the AdaBoost algorithm in terms of its advantages and shortcomings in tacking the class imbalance problem leads to the exploration of three cost-sensitive boosting algorithms, which are developed by introducing cost items into the learning framework of AdaBoost. Further analysis shows that one of the proposed algorithms tallies with the stagewise additive modelling in statistics to minimize the cost exponential loss. These boosting algorithms are also studied with respect to their weighting strategies towards different types of samples, and their effectiveness in identifying rare cases through experiments on several real world medical data sets, where the class imbalance problem prevails. (c) 2007 Pattern Recognition Society. Published by Elsevier Ltd. All rights reserved.