Measuring the Algorithmic Convergence of Randomized Ensembles: The Regression Setting

Measuring the Algorithmic Convergence of Randomized Ensembles: The Regression Setting
复制标题

DOI:
10.1137/20m1343300
复制
发表时间:
2019-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Miles E. Lopes;Suofei Wu;Thomas C.M. Lee
Miles E. Lopes;Suofei Wu;Thomas C.M. Lee
中科院分区:
其他
文献类型:
--
作者:
Miles E. Lopes;Suofei Wu;Thomas C.M. Lee

文献摘要

相似文献

当实现装袋和随机森林等随机集成方法时,会出现一个基本问题:集成是否足够大?特别是,实践者希望严格保证给定的集成的性能几乎与理想的无限集成(在相同的数据上训练)一样好。本文的目的是开发一种引导方法来在回归的背景下解决这个问题——这对我们在分类的背景下的配套论文(Lopes 2019)进行了补充。与分类设置相反,当前的论文表明,可以在更弱的假设下建立所提出的引导程序的理论保证。此外,我们通过展示如何调整该方法来衡量变量选择的算法收敛性来说明该方法的灵活性。最后,我们提供了数值结果,证明该方法在多种情况下都能很好地发挥作用。
When randomized ensemble methods such as bagging and random forests are implemented, a basic question arises: Is the ensemble large enough? In particular, the practitioner desires a rigorous guarantee that a given ensemble will perform nearly as well as an ideal infinite ensemble (trained on the same data). The purpose of the current paper is to develop a bootstrap method for solving this problem in the context of regression --- which complements our companion paper in the context of classification (Lopes 2019). In contrast to the classification setting, the current paper shows that theoretical guarantees for the proposed bootstrap can be established under much weaker assumptions. In addition, we illustrate the flexibility of the method by showing how it can be adapted to measure algorithmic convergence for variable selection. Lastly, we provide numerical results demonstrating that the method works well in a range of situations.