A comparison of decision tree ensemble creation techniques

A comparison of decision tree ensemble creation techniques
复制标题

DOI:
10.1109/tpami.2007.250609
复制
发表时间:
2007-01-01
影响因子:
23.6
通讯作者:
Kegelmeyer, W. P.
Kegelmeyer, W. P.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Banfield, Robert E.;Hall, Lawrence O.;Kegelmeyer, W. P.

文献摘要

被引文献

相似文献

我们通过实验评估包装和其他七种基于随机的方法来创建决策树分类器的集合。对57个公开数据集的实验结果进行了统计测试。当测试了跨验证比较的统计显着性时,最佳方法在统计学上比仅在57个数据集中的八个中袋装更准确。另外,检查整个数据集中算法的平均排名,我们发现升高,随机森林和随机树在统计学上比包装好得多。因为我们的结果表明使用适当的合奏尺寸很重要,所以我们引入了一种算法,该算法决定何时为集合创建了足够数量的分类器。我们的算法使用了止袋外误差估计,并证明可以为那些将袋装纳入整体构造的方法提供准确的合奏。
We experimentally evaluate bagging and seven other randomization-based approaches to creating an ensemble of decision tree classifiers. Statistical tests were performed on experimental results from 57 publicly available data sets. When cross-validation comparisons were tested for statistical significance, the best method was statistically more accurate than bagging on only eight of the 57 data sets. Alternatively, examining the average ranks of the algorithms across the group of data sets, we find that boosting, random forests, and randomized trees are statistically significantly better than bagging. Because our results suggest that using an appropriate ensemble size is important, we introduce an algorithm that decides when a sufficient number of classifiers has been created for an ensemble. Our algorithm uses the out-of-bag error estimate, and is shown to result in an accurate ensemble for those methods that incorporate bagging into the construction of the ensemble.