Scalable and Efficient Hypothesis Testing with Random Forests

Scalable and Efficient Hypothesis Testing with Random Forests
复制标题

DOI:
--
复制
发表时间:
2019-04
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
T. Coleman;Wei Peng;L. Mentch
T. Coleman;Wei Peng;L. Mentch
中科院分区:
其他
文献类型:
--
作者:
T. Coleman;Wei Peng;L. Mentch

文献摘要

相似文献

在过去的十年中,随机森林已经成为最准确和最流行的监督学习方法之一。虽然它们的黑箱性质使它们的数学分析变得困难,但最近的工作已经通过考虑子采样代替自举建立了重要的统计特性,如一致性和渐近正态性。虽然这样的结果打开了传统的推理程序的大门,所有正式的方法,迄今为止提出的地方严格限制的测试框架和他们的计算开销排除其实际的科学用途。在这里,我们提出了一个置换式的测试方法,正式评估功能的重要性。我们建立了渐近有效性的测试,通过交换参数,并表明该测试保持高功率的数量级较少的计算。同样重要的是,该过程很容易扩展到大数据设置,其中可以使用大型训练和测试集,而无需构建额外的模型。随机森林最近表现出希望的生态数据的模拟和应用。
Throughout the last decade, random forests have established themselves as among the most accurate and popular supervised learning methods. While their black-box nature has made their mathematical analysis difficult, recent work has established important statistical properties like consistency and asymptotic normality by considering subsampling in lieu of bootstrapping. Though such results open the door to traditional inference procedures, all formal methods suggested thus far place severe restrictions on the testing framework and their computational overhead precludes their practical scientific use. Here we propose a permutation-style testing approach to formally assess feature significance. We establish asymptotic validity of the test via exchangeability arguments and show that the test maintains high power with orders of magnitude fewer computations. As importantly, the procedure scales easily to big data settings where large training and testing sets may be employed without the need to construct additional models. Simulations and applications to ecological data where random forests have recently shown promise are provided.