The random subspace method for constructing decision forests

The random subspace method for constructing decision forests
复制标题

DOI:
10.1109/34.709601
复制
发表时间:
1998-08-01
影响因子:
23.6
通讯作者:
Ho, TK
Ho, TK
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ho, TK

文献摘要

被引文献

相似文献

以前对决策树的关注主要集中在分裂准则和树大小的优化上。过拟合和达到最大精度之间的矛盾很少得到解决。提出了一种构造基于决策树的分类器的方法,该方法在训练数据上保持最高的精度,并且随着复杂度的增加而提高泛化精度。该分类器由多个树构成,系统地通过伪随机选择的特征向量的分量的子集,也就是说,在随机选择的子空间中构建的树。通过在公开数据集上的实验,将子空间方法与单树分类器和其他森林构造方法进行了比较,证明了该方法的优越性。我们还讨论了森林中树木之间的独立性,并将其与组合分类精度相关联。
Much of previous attention on decision trees focuses on the splitting criteria and optimization of tree sizes. The dilemma between overfitting and achieving maximum accuracy is seldom resolved. A method to construct a decision tree based classifier is proposed that maintains highest accuracy on training data and improves on generalization accuracy as it grows in complexity. The classifier consists of multiple trees constructed systematically by pseudorandomly selecting subsets of components of the feature vector, that is, trees constructed in randomly chosen subspaces. The subspace method is compared to single-tree classifiers and other forest construction methods by experiments on publicly available datasets, where the method's superiority is demonstrated. We also discuss independence between trees in a forest and relate that to the combined classification accuracy.