The Effect of Splitting on Random Forests.

The Effect of Splitting on Random Forests.
复制标题

DOI:
10.1007/s10994-014-5451-2
复制
发表时间:
2015-04
期刊:
影响因子:
7.5
通讯作者:
Ishwaran H
Ishwaran H
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ishwaran H

文献摘要

参考文献

被引文献

相似文献

针对回归和分类问题,系统地研究了分裂规则对随机森林的影响。详细研究了一类加权分裂规则,它包括CART加权方差分裂和基尼指数分裂作为特例,并证明了它对信号和噪声具有独特的自适应特性。对于有噪声的变量,我们证明加权分裂有利于端切分裂。虽然传统上认为端切劈裂对于单株树木是不可取的,但我们认为对于生长较深的树木(RF的商标),端切劈裂是有用的,因为:(A)它使样本大小最大化,使得树木有可能从糟糕的劈裂中恢复,以及(B)如果树枝因噪声而反复分裂,则将达到树的最小节点大小,从而促进不良枝条的终止。对于强变量,加权方差分裂在目标函数的曲率点处具有理想的分裂性质。对于未加权和重加权分割规则,这种对噪声和信号的适应性都不成立。后一种规则要么过于贪婪,不善于识别嘈杂场景,要么过于咄咄逼人,不善于识别信号。这些结果也揭示了纯粹的随机分裂,并表明这种规则是最不有效的。另一方面,由于随机化规则的计算效率是可取的,我们引入了一种采用随机分裂点选择的混合方法,该方法在保持计算效率的同时保持了加权分裂规则的自适应特性。
The effect of a splitting rule on random forests (RF) is systematically studied for regression and classification problems. A class of weighted splitting rules, which includes as special cases CART weighted variance splitting and Gini index splitting, are studied in detail and shown to possess a unique adaptive property to signal and noise. We show for noisy variables that weighted splitting favors end-cut splits. While end-cut splits have traditionally been viewed as undesirable for single trees, we argue for deeply grown trees (a trademark of RF) end-cut splitting is useful because: (a) it maximizes the sample size making it possible for a tree to recover from a bad split, and (b) if a branch repeatedly splits on noise, the tree minimal node size will be reached which promotes termination of the bad branch. For strong variables, weighted variance splitting is shown to possess the desirable property of splitting at points of curvature of the underlying target function. This adaptivity to both noise and signal does not hold for unweighted and heavy weighted splitting rules. These latter rules are either too greedy, making them poor at recognizing noisy scenarios, or they are overly ECP aggressive, making them poor at recognizing signal. These results also shed light on pure random splitting and show that such rules are the least effective. On the other hand, because randomized rules are desirable because of their computational efficiency, we introduce a hybrid method employing random split-point selection which retains the adaptive property of weighted splitting rules while remaining computational efficient.
DOI: 10.1214/aos/1176347498
发表时间: 1990-03-01
影响因子: 4.5
作者:
KIM, JY;POLLARD, D
通讯作者: POLLARD, D
DOI: 10.1023/a:1007607513941
发表时间: 2000-08-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Dietterich, TG
通讯作者: Dietterich, TG
DOI: 10.1023/a:1018094028462
发表时间: 1996-07-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Breiman, L
通讯作者: Breiman, L
DOI: 10.1007/s10994-006-6226-1
发表时间: 2006-04-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Geurts, P;Ernst, D;Wehenkel, L
通讯作者: Wehenkel, L
DOI: 10.1080/10485252.2012.677843
发表时间: 2012-01-01
影响因子: 1.2
作者:
Genuer, Robin
通讯作者: Genuer, Robin