Boosting Random Forests to Reduce Bias; One-Step Boosted Forest and Its Variance Estimate

Boosting Random Forests to Reduce Bias; One-Step Boosted Forest and Its Variance Estimate
复制标题

DOI:
10.1080/10618600.2020.1820345
复制
发表时间:
2018-03
影响因子:
2.4
通讯作者:
Indrayudh Ghosal;G. Hooker
Indrayudh Ghosal;G. Hooker
中科院分区:
数学2区
文献类型:
--
作者:
Indrayudh Ghosal;G. Hooker

文献摘要

被引文献

相似文献

摘要在本文中,我们提出了在回归设置中使用增强原理来降低随机森林预测的偏差。从原始随机森林拟合中提取残差,然后对这些残差进行另一个随机森林拟合。我们称这两个随机森林的和为一步推进森林。我们用模拟和真实数据表明,一步增强森林与原始随机森林相比具有更小的偏差。本文还利用推广的无限小折刀估计,给出了一步推进林的方差估计。利用该方差估计,我们可以构建增强森林的预测区间,并证明它们具有良好的覆盖概率。结合偏差减少和方差估计,我们发现一步增强森林的预测均方误差显著降低,从而提高了预测性能。当应用于UCI数据库的数据集时,一步增强森林算法的性能优于随机森林和梯度增强机算法。从理论上讲,我们也可以将这样的提升过程扩展到多个步骤,并且本文中概述的相同原则可以用于查找此类预测器的方差估计。这样的提升将进一步减少偏差,但它有过拟合的风险,也增加了计算负担。本文的补充材料可在网上获得。
Abstract In this article, we propose using the principle of boosting to reduce the bias of a random forest prediction in the regression setting. From the original random forest fit, we extract the residuals and then fit another random forest to these residuals. We call the sum of these two random forests a one-step boosted forest. We show with simulated and real data that the one-step boosted forest has a reduced bias compared to the original random forest. The article also provides a variance estimate of the one-step boosted forest by an extension of the infinitesimal Jackknife estimator. Using this variance estimate, we can construct prediction intervals for the boosted forest and we show that they have good coverage probabilities. Combining the bias reduction and the variance estimate, we show that the one-step boosted forest has a significant reduction in predictive mean squared error and thus an improvement in predictive performance. When applied on datasets from the UCI database, one-step boosted forest performs better than random forest and gradient boosting machine algorithms. Theoretically, we can also extend such a boosting process to more than one step and the same principles outlined in this article can be used to find variance estimates for such predictors. Such boosting will reduce bias even further but it risks over-fitting and also increases the computational burden. Supplementary materials for this article are available online.