Highly scalable and robust rule learner: Performance evaluation and comparison

Highly scalable and robust rule learner: Performance evaluation and comparison
复制标题

DOI:
10.1109/tsmcb.2005.852983
复制
发表时间:
2006-02-01
影响因子:
--
通讯作者:
Dick, S
Dick, S
中科院分区:
其他
文献类型:
--
作者:
Kurgan, LA;Cios, KJ;Dick, S

文献摘要

被引文献

相似文献

商业智能和生物信息学应用越来越需要挖掘由数百万个数据点组成的数据集,或者为大公司和制药公司制作实时企业级决策支持系统。在所有情况下,都需要一个底层数据挖掘系统,并且这个挖掘系统必须具有高度可扩展性。为此,我们描述了一种新的规则学习器,称为DataSqueezer。该学习算法属于归纳监督规则提取算法。DataSqueezer是一个简单的、贪婪的规则构建器,它从标记的输入数据生成一组生产规则。尽管相对简单,DataSqueezer是一个非常有效的学习工具。该算法生成的规则紧凑、易于理解,并且与其他最先进的规则提取算法生成的规则具有相当的准确性。DataSqueezer的主要优点是非常高的效率,并且可以抵抗丢失数据。DataSqueezer随着训练示例的数量呈现对数线性渐近复杂性,并且比其他最先进的规则学习器更快。通过与其他学习器的广泛实验比较,该学习器对大量缺失数据也具有鲁棒性。因此,DataSqueezer非常适合现代数据挖掘和商业智能任务,这些任务通常涉及具有大量缺失数据的庞大数据集。
Business intelligence and bioinformatics applications increasingly require the mining of datasets consisting of millions of data points, or crafting real-time enterprise-level decision support systems for large corporations and drug companies. In all cases, there needs to be an underlying data mining system, and this mining system must be highly scalable. To this end, we describe a new rule learner called DataSqueezer. The learner belongs to the family of inductive supervised rule extraction algorithms. DataSqueezer is a simple, greedy, rule builder that generates a set of production rules from labeled input data. In spite of its relative simplicity, DataSqueezer is a very effective learner. The rules generated by the algorithm are compact, comprehensible, and have accuracy comparable to rules generated by other state-of-the-art rule extraction algorithms. The main advantages of DataSqueezer are very high efficiency, and missing data resistance. DataSqueezer exhibits log-linear asymptotic complexity with the number of training examples, and it is faster than other state-of-the-art rule learners. The learner is also robust to large quantities of missing data, as verified by extensive experimental comparison with the other learners. DataSqueezer is thus well suited to modern data mining and business intelligence tasks, which commonly involve huge datasets with a large fraction of missing data.