Randomization as Regularization: A Degrees of Freedom Explanation for Random Forest Success

Randomization as Regularization: A Degrees of Freedom Explanation for Random Forest Success
复制标题

DOI:
--
复制
发表时间:
2019-11
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
L. Mentch;Siyu Zhou
L. Mentch;Siyu Zhou
中科院分区:
其他
文献类型:
--
作者:
L. Mentch;Siyu Zhou

文献摘要

被引文献

相似文献

随机森林仍然是最受欢迎的现成监督机器学习工具之一,在回归和分类设置中具有良好的预测准确性记录。尽管它们在经验上取得了成功,最近也有一批研究它们统计特性的工作,但对它们的成功还没有提出一个完整而令人满意的解释。在这里,我们的目标是在这个方向上迈出一步,通过证明注入到单个树中的额外随机性作为一种形式的隐式正则化,使随机森林成为低信噪比(SNR)设置中的理想模型。具体来说,从模型复杂性的角度来看,我们表明,mtry参数在随机森林中的收缩惩罚在显式正则化回归过程,如套索和岭回归相同的目的。为了突出这一点,我们设计了一个随机线性模型为基础的前向选择过程,旨在作为一个类似的基于树的随机森林,并证明其令人惊讶的强大的经验表现。提供了许多关于真实的和合成数据的演示。
Random forests remain among the most popular off-the-shelf supervised machine learning tools with a well-established track record of predictive accuracy in both regression and classification settings. Despite their empirical success as well as a bevy of recent work investigating their statistical properties, a full and satisfying explanation for their success has yet to be put forth. Here we aim to take a step forward in this direction by demonstrating that the additional randomness injected into individual trees serves as a form of implicit regularization, making random forests an ideal model in low signal-to-noise ratio (SNR) settings. Specifically, from a model-complexity perspective, we show that the mtry parameter in random forests serves much the same purpose as the shrinkage penalty in explicitly regularized regression procedures like lasso and ridge regression. To highlight this point, we design a randomized linear-model-based forward selection procedure intended as an analogue to tree-based random forests and demonstrate its surprisingly strong empirical performance. Numerous demonstrations on both real and synthetic data are provided.