Addressing imbalance in multilabel classification: Measures and random resampling algorithms

Addressing imbalance in multilabel classification: Measures and random resampling algorithms
复制标题

DOI:
10.1016/j.neucom.2014.08.091
复制
发表时间:
2015-09-02
期刊:
影响因子:
6
通讯作者:
Herrera, Francisco
Herrera, Francisco
中科院分区:
计算机科学2区
文献类型:
--
作者:
Charte, Francisco;Rivera, Antonio J.;Herrera, Francisco

文献摘要

被引文献

相似文献

本文的目的是分析多标签场景中的不平衡学习任务,旨在实现两个不同的目标。第一个是提出专门的措施,旨在评估多标签数据集(MLDs)的不平衡水平。使用这些措施,我们将能够得出结论,MLD是不平衡的,因此将需要一个适当的治疗。第二个目标是提出几种算法,旨在减少不平衡的MLD在一个独立的分类器的方式,通过resternation技术。研究了两种不同的方法来划分少数群体和多数群体中的实例。其中一个将每个标签组合视为类标识符,而另一个对每个标签不平衡水平进行单独评估。针对每种方法提出了一种随机欠采样和随机过采样算法,并给出了四种不同的算法。所有这些都是实验测试和他们的有效性进行了统计评估。从所获得的结果,一组指导方针,以显示这些方法应该应用时,也提供了。(C)2015 Elsevier B.V.版权所有。
The purpose of this paper is to analyze the imbalanced learning task in the multilabel scenario, aiming to accomplish two different goals. The first one is to present specialized measures directed to assess the imbalance level in multilabel datasets (MLDs). Using these measures we will be able to conclude which MLDs are imbalanced, and therefore would need an appropriate treatment The second objective is to propose several algorithms designed to reduce the imbalance in MLDs in a classifier-independent way, by means of resampling techniques. Two different approaches to divide the instances in minority and majority groups are studied. One of them considers each label combination as class identifier, whereas the other one performs an individual evaluation of each label imbalance level. A random undersampling and a random oversampling algorithm are proposed for each approach, giving as result four different algorithms. All of them are experimentally tested and their effectiveness is statistically evaluated. From the results obtained, a set of guidelines directed to show when these methods should be applied is also provided. (C) 2015 Elsevier B.V. All rights reserved.