On Mining Instance-Centric Classification Rules

On Mining Instance-Centric Classification Rules
复制标题

DOI:
10.1109/tkde.2006.179
复制
发表时间:
2006-11
影响因子:
8.9
通讯作者:
Jianyong Wang;G. Karypis
Jianyong Wang;G. Karypis
中科院分区:
计算机科学2区
文献类型:
--
作者:
Jianyong Wang;G. Karypis

文献摘要

被引文献

相似文献

许多研究表明,基于规则的分类器在分类和稀疏高维数据库中表现良好。然而,许多基于规则的分类器的一个基本限制是,它们通过采用各种启发式方法来修剪搜索空间并基于顺序数据库覆盖范式选择规则来发现规则。因此,他们使用的最终规则集可能不是训练数据库中某些实例的全局最佳规则。更糟糕的是,这些算法未能充分利用一些更有效的搜索空间修剪方法,以扩展到大型数据库。在本文中,我们提出了一个新的分类器,和谐,它直接挖掘最终的分类规则集。HARMONY使用以实例为中心的规则生成方法,它可以确保对于每个训练实例,覆盖该实例的最高置信度规则之一包含在最终规则集中,这有助于提高分类器的整体准确性。通过在规则发现过程中引入多种新颖的搜索策略和剪枝方法,HARMONY也具有较高的效率和良好的可扩展性。我们对一些大型文本和分类数据库进行了全面的性能研究,结果表明,HARMONY在准确性和计算效率方面优于许多知名的分类器,并且在数据库大小方面具有很好的扩展性
Many studies have shown that rule-based classifiers perform well in classifying categorical and sparse high-dimensional databases. However, a fundamental limitation with many rule-based classifiers is that they find the rules by employing various heuristic methods to prune the search space and select the rules based on the sequential database covering paradigm. As a result, the final set of rules that they use may not be the globally best rules for some instances in the training database. To make matters worse, these algorithms fail to fully exploit some more effective search space pruning methods in order to scale to large databases. In this paper, we present a new classifier, HARMONY, which directly mines the final set of classification rules. HARMONY uses an instance-centric rule-generation approach and it can assure that, for each training instance, one of the highest-confidence rules covering this instance is included in the final rule set, which helps in improving the overall accuracy of the classifier. By introducing several novel search strategies and pruning methods into the rule discovery process, HARMONY also has high efficiency and good scalability. Our thorough performance study with some large text and categorical databases has shown that HARMONY outperforms many well-known classifiers in terms of both accuracy and computational efficiency and scales well with regard to the database size