Online sparse class imbalance learning on big data

Online sparse class imbalance learning on big data
复制标题

DOI:
10.1016/j.neucom.2016.07.040
复制
发表时间:
2016-12
期刊:
影响因子:
6
通讯作者:
Chandresh Kumar Maurya;Durga Toshniwal;G. V. Venkoparao
Chandresh Kumar Maurya;Durga Toshniwal;G. V. Venkoparao
中科院分区:
计算机科学2区
文献类型:
--
作者:
Chandresh Kumar Maurya;Durga Toshniwal;G. V. Venkoparao

文献摘要

被引文献

相似文献

班级不平衡学习是对某些班级比其他班级出现频率更高的问题的研究。大多数现有的研究这个问题的作品假设数据集是密集的,并没有利用丰富的数据结构。一个这样的结构是稀疏性。在目前的工作中,我们专注于解决稀疏性假设下的类不平衡问题。更具体地说,一个众所周知的Gmeanmetric类不平衡学习问题的二进制分类设置已被最大化,这导致在一个非凸的损失函数。利用凸松弛技术将非凸问题转化为凸问题。在目前的工作中的问题制定使用L1正则化邻近学习框架,并通过加速随机邻近梯度下降算法求解。我们在本文中的目的是显示:(i)应用近似算法来解决真实的世界问题(类不平衡);(ii)它如何扩展到大数据;以及(iii)它如何在几个基准数据集上的Gmean,F-measure和Mistake rateo方面优于最近提出的一些算法。
Class imbalance learning is the study of problems in which some classes appear more frequently than the others. Most existing works that study this problem assume data set to be dense and do not exploit the rich structure of the data. One such structure is the sparsity. In the present work, we focus on solving the class imbalance problem under the sparsity assumption. More specifically, a well-knownGmeanmetric for class imbalance learning problem in binary classification setting has been maximized, which results in a non-convex loss function. Convex relaxation techniques are used to convert the non-convex problem to the convex problem. The problem formulation in the present work usesL1regularized proximal learning framework and is solved via accelerated-stochastic-proximal gradient descent algorithm. Our aim in the paper is to show: (i) the application of proximal algorithms to solve real world problems (class imbalance); (ii) how it scales to Big data; and (iii) how it outperforms some recently proposed algorithms in terms ofGmean, F-measureandMistake rateon several benchmark data sets.