A novel evolutionary data mining algorithm with applications to churn prediction

A novel evolutionary data mining algorithm with applications to churn prediction
复制标题

DOI:
10.1109/tevc.2003.819264
复制
发表时间:
2003-12
期刊:
IEEE Trans. Evol. Comput.
影响因子:
--
通讯作者:
Wai-Ho Au;Keith C. C. Chan-Keith-C.-C.-Chan-145003399;X. Yao
Wai-Ho Au;Keith C. C. Chan-Keith-C.-C.-Chan-145003399;X. Yao
中科院分区:
其他
文献类型:
--
作者:
Wai-Ho Au;Keith C. C. Chan-Keith-C.-C.-Chan-145003399;X. Yao

文献摘要

被引文献

相似文献

分类是数据挖掘研究中的一个重要课题。给定一组数据记录,其中每一个属于一个预定义的类,分类问题涉及的分类规则的发现,可以允许记录与未知的类成员被正确分类。已经开发了许多算法来挖掘大数据集的分类模型,它们已被证明是非常有效的。然而,当涉及到确定每个分类的可能性时,其中许多分类的设计并没有考虑到这种目的。为此,它们不容易适用于诸如流失预测之类的问题。对于这样的应用,目标不仅是预测用户是否会从一个运营商切换到另一个运营商,预测用户这样做的可能性也很重要。这样做的原因是,运营商可以选择向那些被预测具有较高流失可能性的订户提供特殊的个性化报价和服务。鉴于其重要性,我们提出了一种新的数据挖掘算法,称为数据挖掘的进化学习(DMEL),处理分类问题,其中每个预测的准确性进行了估计。在执行其任务时,DMEL使用具有以下特征的进化方法来搜索可能的规则空间:1)进化过程开始于生成一阶规则的初始集合(即,具有一个合取/条件的规则),并基于这些规则,(两个或多个合取)被迭代地获得; 2)当识别感兴趣的规则时,使用客观的兴趣度度量; 3)染色体的适应度被定义为使用其编码的规则可以正确地确定记录的属性值的概率;以及4)估计所作出的预测(或分类)的可能性,使得可以根据订户流失的可能性对订户进行排序。不同数据集的实验表明,DMEL能够有效地发现有趣的分类规则。特别地,当应用于真实的电信用户数据时,它能够在不同的流失率下准确地预测流失。
Classification is an important topic in data mining research. Given a set of data records, each of which belongs to one of a number of predefined classes, the classification problem is concerned with the discovery of classification rules that can allow records with unknown class membership to be correctly classified. Many algorithms have been developed to mine large data sets for classification models and they have been shown to be very effective. However, when it comes to determining the likelihood of each classification made, many of them are not designed with such purpose in mind. For this, they are not readily applicable to such problems as churn prediction. For such an application, the goal is not only to predict whether or not a subscriber would switch from one carrier to another, it is also important that the likelihood of the subscriber's doing so be predicted. The reason for this is that a carrier can then choose to provide a special personalized offer and services to those subscribers who are predicted with higher likelihood to churn. Given its importance, we propose a new data mining algorithm, called data mining by evolutionary learning (DMEL), to handle classification problems of which the accuracy of each predictions made has to be estimated. In performing its tasks, DMEL searches through the possible rule space using an evolutionary approach that has the following characteristics: 1) the evolutionary process begins with the generation of an initial set of first-order rules (i.e., rules with one conjunct/condition) using a probabilistic induction technique and based on these rules, rules of higher order (two or more conjuncts) are obtained iteratively; 2) when identifying interesting rules, an objective interestingness measure is used; 3) the fitness of a chromosome is defined in terms of the probability that the attribute values of a record can be correctly determined using the rules it encodes; and 4) the likelihood of predictions (or classifications) made are estimated so that subscribers can be ranked according to their likelihood to churn. Experiments with different data sets showed that DMEL is able to effectively discover interesting classification rules. In particular, it is able to predict churn accurately under different churn rates when applied to real telecom subscriber data.