Generating Accurate Rule Sets Without Global Optimization

Generating Accurate Rule Sets Without Global Optimization
复制标题

DOI:
--
复制
发表时间:
1998-07
期刊:
--
影响因子:
--
通讯作者:
E. Frank;I. Witten
E. Frank;I. Witten
中科院分区:
其他
文献类型:
--
作者:
E. Frank;I. Witten

文献摘要

被引文献

相似文献

规则学习的两个主要方案,C4.5和RIPPER,都分为两个阶段。首先,他们归纳出一个初始规则集,然后使用一个相当复杂的优化阶段对其进行优化,该阶段丢弃(C4.5)或调整(RIPPER)单个规则,以使它们更好地协同工作。相比之下,本文展示了如何一次一个规则地学习好的规则集,而不需要任何全局优化。我们提出了一种通过重复生成部分决策树来推断规则的算法,从而结合了规则生成的两种主要范式--从决策树创建规则和分离并征服规则学习技术。该算法简单而优雅:尽管如此,在标准数据集上的实验表明,它产生的规则集与C4.5产生的规则集一样准确,大小相似,比RIPPER更准确。此外,它有效地运行,并且因为它避免了后处理,所以不会遭受C4.5方法被批评的病理性示例集的极慢性能。
The two dominant schemes for rule-learning, C4.5 and RIPPER, both operate in two stages. First they induce an initial rule set and then they refine it using a rather complex optimization stage that discards (C4.5) or adjusts (RIPPER) individual rules to make them work better together. In contrast, this paper shows how good rule sets can be learned one rule at a time, without any need for global optimization. We present an algorithm for inferring rules by repeatedly generating partial decision trees, thus combining the two major paradigms for rule generation—creating rules from decision trees and the separate-and-conquer rule-learning technique. The algorithm is straightforward and elegant: despite this, experiments on standard datasets show that it produces rule sets that are as accurate as and of similar size to those generated by C4.5, and more accurate than RIPPER’s. Moreover, it operates efficiently, and because it avoids postprocessing, does not suffer the extremely slow performance on pathological example sets for which the C4.5 method has been criticized.