When do random forests fail?

When do random forests fail?
复制标题

DOI:
--
复制
发表时间:
2018-12
期刊:
--
影响因子:
--
通讯作者:
Cheng Tang;D. Garreau;U. V. Luxburg
Cheng Tang;D. Garreau;U. V. Luxburg
中科院分区:
其他
文献类型:
--
作者:
Cheng Tang;D. Garreau;U. V. Luxburg

文献摘要

被引文献

相似文献

随机森林是一种学习算法,它构建了大量的随机树集合,并通过对单个树的预测进行平均来进行预测。在本文中,我们考虑了各种树的结构,并研究如何选择的参数影响的泛化误差的随机森林的样本大小趋于无穷大。我们表明,在树的建设阶段的数据点的二次采样是很重要的:森林可以变得不一致,无论是没有二次采样或过于严重的二次采样。因此,如果不使用子采样,即使高度随机化的树也会导致不一致的森林,这意味着随机森林的一些常用设置可能不一致。作为第二个结果,我们可以证明,在最近邻搜索中具有良好性能的树可能是随机森林的糟糕选择。
Random forests are learning algorithms that build large collections of random trees and make predictions by averaging the individual tree predictions. In this paper, we consider various tree constructions and examine how the choice of parameters affects the generalization error of the resulting random forests as the sample size goes to infinity. We show that subsampling of data points during the tree construction phase is important: Forests can become inconsistent with either no subsampling or too severe subsampling. As a consequence, even highly randomized trees can lead to inconsistent forests if no subsampling is used, which implies that some of the commonly used setups for random forests can be inconsistent. As a second consequence we can show that trees that have good performance in nearest-neighbor search can be a poor choice for random forests.