The application of data mining techniques in financial fraud detection: A classification framework and an academic review of literature

The application of data mining techniques in financial fraud detection: A classification framework and an academic review of literature
复制标题

DOI:
10.1016/j.dss.2010.08.006
复制
发表时间:
2011-02-01
影响因子:
7.5
通讯作者:
Sun, Xin
Sun, Xin
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ngai, E. W. T.;Hu, Yong;Sun, Xin

文献摘要

被引文献

相似文献

本文对数据挖掘技术在金融欺诈检测中的应用进行了综述和分类。虽然金融欺诈检测(FFD)是一个非常重要的新兴课题,一个全面的文献综述的主题还没有进行。因此,本文代表了第一个系统的,可识别的和全面的学术文献综述的数据挖掘技术,已应用于FFD。分析了1997年至2008年期间发表的49篇关于该主题的期刊文章,并将其分为四类金融欺诈(银行欺诈,保险欺诈,证券和商品欺诈以及其他相关的金融欺诈)和六类数据挖掘技术(分类,回归,聚类,预测,异常值检测和可视化)。这项审查的结果清楚地表明,数据挖掘技术已被广泛应用于检测保险欺诈,虽然公司欺诈和信用卡欺诈也吸引了大量的关注,在最近几年。相比之下,我们发现明显缺乏对抵押贷款欺诈,洗钱,证券和商品欺诈的研究。用于FFD的主要数据挖掘技术是逻辑模型,神经网络,贝叶斯信念网络和决策树,所有这些都为欺诈数据的检测和分类中固有的问题提供了主要的解决方案。本文还讨论了FFD与行业需求之间的差距,以鼓励对被忽视的主题进行更多的研究,并对进一步的FFD研究提出了几项建议。皇冠版权所有(C)2010由爱思唯尔B.V.出版保留所有权利。
This paper presents a review of and classification scheme for - the literature on the application of data mining techniques for the detection of financial fraud. Although financial fraud detection (FFD) is an emerging topic of great importance, a comprehensive literature review of the subject has yet to be carried out. This paper thus represents the first systematic, identifiable and comprehensive academic literature review of the data mining techniques that have been applied to FFD. 49 journal articles on the subject published between 1997 and 2008 was analyzed and classified into four categories of financial fraud (bank fraud, insurance fraud, securities and commodities fraud, and other related financial fraud) and six classes of data mining techniques (classification, regression, clustering, prediction, outlier detection, and visualization). The findings of this review clearly show that data mining techniques have been applied most extensively to the detection of insurance fraud, although corporate fraud and credit card fraud have also attracted a great deal of attention in recent years. In contrast, we find a distinct lack of research on mortgage fraud, money laundering, and securities and commodities fraud. The main data mining techniques used for FFD are logistic models, neural networks, the Bayesian belief network, and decision trees, all of which provide primary solutions to the problems inherent in the detection and classification of fraudulent data. This paper also addresses the gaps between FFD and the needs of the industry to encourage additional research on neglected topics, and concludes with several suggestions for further FFD research. Crown Copyright (C) 2010 Published by Elsevier B.V. All rights reserved.