A Survey of Predictive Modeling on Im balanced Domains

A Survey of Predictive Modeling on Im balanced Domains
复制标题

DOI:
10.1145/2907070
复制
发表时间:
2016-11-01
影响因子:
16.6
通讯作者:
Ribeiro, Rita P.
Ribeiro, Rita P.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Branco, Paula;Torgo, Luis;Ribeiro, Rita P.

文献摘要

被引文献

相似文献

许多现实世界的数据挖掘应用涉及使用具有目标变量的强烈不平衡分布的数据集来获得预测模型。通常,该目标变量的最小公共值与对于最终用户高度相关的事件相关联(例如,欺诈检测、股票市场的异常回报、对灾难的预期等)。此外,这些事件可能具有不同的成本和收益,当与其中一些事件在可用的训练数据上的稀有性相关联时,会给预测建模技术带来严重的问题。本文介绍了处理预测分析的这些重要应用的现有技术的调查。虽然大多数现有的工作地址分类任务(名义上的目标变量),我们还描述了设计来处理类似的问题,在回归任务(数字目标变量)的方法。在这项调查中,我们讨论了不平衡领域提出的主要挑战,提出了问题的定义,描述了这些任务的主要方法,提出了方法的分类,总结了现有的比较研究的结论,以及一些方法的理论分析,并参考预测建模中的一些相关问题。
Many real-world data-mining applications involve obtaining predictive models using datasets with strongly imbalanced distributions of the target variable. Frequently, the least-common values of this target variable are associated with events that are highly relevant for end users (e.g., fraud detection, unusual returns on stock markets, anticipation of catastrophes, etc.). Moreover, the events may have different costs and benefits, which, when associated with the rarity of some of them on the available training data, creates serious problems to predictive modeling techniques. This article presents a survey of existing techniques for handling these important applications of predictive analytics. Although most of the existing work addresses classification tasks (nominal target variables), we also describe methods designed to handle similar problems within regression tasks (numeric target variables). In this survey, we discuss the main challenges raised by imbalanced domains, propose a definition of the problem, describe the main approaches to these tasks, propose a taxonomy of the methods, summarize the conclusions of existing comparative studies as well as some theoretical analyses of some methods, and refer to some related problems within predictive modeling.