Classification cost: An empirical comparison among traditional classifier, Cost-Sensitive Classifier, and MetaCost

Classification cost: An empirical comparison among traditional classifier, Cost-Sensitive Classifier, and MetaCost
复制标题

DOI:
10.1016/j.eswa.2011.09.071
复制
发表时间:
2012-03
期刊:
Expert Syst. Appl.
影响因子:
--
通讯作者:
Jungeun Kim;Keunho Choi;Gunwoo Kim;Yongmoo Suh
Jungeun Kim;Keunho Choi;Gunwoo Kim;Yongmoo Suh
中科院分区:
其他
文献类型:
--
作者:
Jungeun Kim;Keunho Choi;Gunwoo Kim;Yongmoo Suh

文献摘要

被引文献

相似文献

贷款欺诈是金融机构破产的一个关键因素,因此公司努力通过建立主动欺诈预测模型来减少欺诈造成的损失。然而,欺诈检测仍然存在两个关键问题需要解决:(1)大多数预测模型中I型错误和II型错误之间缺乏成本敏感性,以及(2)由于稀疏的欺诈相关数据,用于欺诈检测的数据集中的类分布高度偏斜。本文的目的是研究分类成本是否同时受到成本敏感方法和类的偏态分布的影响。为此,我们比较了传统的成本不敏感的分类方法和两个成本敏感的分类方法,成本敏感的分类器(CSC)和MetaCost的分类成本。实验进行了一个主要的金融机构在韩国的信用贷款数据集,而不同的类在数据集中的分布和输入变量的数量。实验表明,当使用MetaCost方法时,以及当非欺诈数据和欺诈数据平衡时,分类成本最低。此外,包括所有拖欠变量的数据集被证明是最有效的降低分类成本。
Loan fraud is a critical factor in the insolvency of financial institutions, so companies make an effort to reduce the loss from fraud by building a model for proactive fraud prediction. However, there are still two critical problems to be resolved for the fraud detection: (1) the lack of cost sensitivity between type I error and type II error in most prediction models, and (2) highly skewed distribution of class in the dataset used for fraud detection because of sparse fraud-related data. The objective of this paper is to examine whether classification cost is affected both by the cost-sensitive approach and by skewed distribution of class. To that end, we compare the classification cost incurred by a traditional cost-insensitive classification approach and two cost-sensitive classification approaches, Cost-Sensitive Classifier (CSC) and MetaCost. Experiments were conducted with a credit loan dataset from a major financial institution in Korea, while varying the distribution of class in the dataset and the number of input variables. The experiments showed that the lowest classification cost was incurred when the MetaCost approach was used and when non-fraud data and fraud data were balanced. In addition, the dataset that includes all delinquency variables was shown to be most effective on reducing the classification cost.