Fast Stochastic Recursive Momentum Methods for Imbalanced Data Mining

Fast Stochastic Recursive Momentum Methods for Imbalanced Data Mining
复制标题

DOI:
10.1109/icdm54844.2022.00068
复制
发表时间:
2022-11
期刊:
2022 IEEE International Conference on Data Mining (ICDM)
影响因子:
--
通讯作者:
Xidong Wu;Feihu Huang;Heng Huang
Xidong Wu;Feihu Huang;Heng Huang
中科院分区:
其他
文献类型:
--
作者:
Xidong Wu;Feihu Huang;Heng Huang

文献摘要

相似文献

标准深度学习模型主要是为平衡数据挖掘任务而设计的,并使用准确率来评估分类器。然而,在许多实际应用中,数据分布是倾斜的。如果将旨在优化准确率的标准模型应用于不平衡数据,由于模型偏向多数类,预测性能可能会很差。为了解决不平衡数据挖掘问题,精确率 - 召回率曲线下面积(AUPRC)被提议作为评估预测模型在不平衡数据集上性能的良好指标,并且在识别具有高预测能力的模型方面表现出出色的能力。为了提高模型的性能,研究人员最近设计了一些方法来直接针对不平衡数据挖掘优化AUPRC。然而,这些方法存在较高的迭代复杂度,因此需要更高效的方法。在本文中,我们提出了一种更快的随机方法(即ROAP),用于基于基于动量的方差减少技术最大化AURPC。我们的新方法基于非参数平均精度(AP)的最大化,AP是AUPRC的一种流行的无偏点估计量,并且本文中的优化目标可以转换为相关复合函数的总和,其中内层函数依赖于内层和外层的随机变量。与先前的方法相比,我们的ROAP算法在寻找一个ϵ - 平稳解时可以实现更低的迭代复杂度$O(\epsilon^{-3})$。此外,我们将我们的方法扩展为一个自适应版本(即AROAP),其具有相同的迭代复杂度$O(\epsilon^{-3})$。据我们所知,本文是第一项表明方差减少方法可以被纳入到最大化AURPC中,以便在不平衡数据集上进行高效数据挖掘的工作。最后,我们使用不同的模型在各种不平衡数据集上进行了广泛的实验,以证明我们新算法的效率。
Standard deep learning models have been mainly designed for balanced data mining tasks and use accuracy to evaluate the classifier. However, in many real-world applications, the distribution of data is skewed. If the standard models, which are designed to optimize the accuracy, are applied to the imbalanced data, the prediction performance could be poor because the model bias towards the majority class. To address the imbalanced data mining problem, areas under precision-recall curves (AUPRC) was proposed as a good measure to evaluate the performance of prediction models on imbalanced data sets, and shows excellent capability in identifying the models with high predictive power. To improve the performance of models, researchers recently design methods to directly optimize AUPRC for imbalanced data mining. However, these approaches suffer from a high iteration complexity and efficient methods are desired. In this paper, we propose a faster stochastic method (i.e., ROAP) for maximizing the AURPC based on the momentum-based variance reduced technique. Our new method is based on the maximization of non-parametric averaged precision (AP), which is a popular unbiased point estimator of AUPRC, and the optimization objective in this paper can be converted into a sum of dependent compositional functions, where the inner functions rely on random variables of both inner and outer levels. Compared to previous methods, our ROAP algorithm can achieve a lower iteration complexity of $O(\epsilon^{-3})$ for finding an ϵ-stationary solution. Furthermore, we extend our method to an adaptive version (i.e., AROAP) with the same iteration complexity of $O(\epsilon^{-3})$. To the best of our knowledge, this paper is the first work showing that the variance reduction method can be incorporated into maximizing the AURPC for efficient data mining on imbalanced datasets. Finally, we conduct extensive experiments on various imbalanced data sets with different models to demonstrate the efficiency of our new algorithms.