Case-Specific Random Forests

Case-Specific Random Forests
复制标题

特定案例的随机森林

DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
D. Nordman
D. Nordman
中科院分区:
--
文献类型:
--
作者:
Ruo Xu;D. Nettleton;D. Nordman

文献摘要

被引文献

相似文献

随机森林(RF)方法是一种用于预测问题的非参数方法。使用RF的标准方法包括生成全局RF以预测所有感兴趣的测试用例。在这篇文章中,我们提出了不同的RF特定于不同的测试用例,即案例特定的随机森林(CSRFs)。与标准RF构建中的装袋过程相反,CSRF算法采用加权自举重新采样来创建单独的树,在该树中,我们将大的权重分配给与先验测试用例非常接近的训练用例。调优方法进行了讨论,以避免过拟合问题。模拟和真实的数据实例表明,CSRF构造中使用的加权Bootstrap方法可以改善特定情况下的预测。我们还提出了一个新的情况下,特定的变量重要性(CSVI)的措施,作为一种方式来比较相对预测变量的重要性,预测一个特定的情况。建立特定于案例的预测器的想法可能可以推广到其他领域。
Random forest (RF) methodology is a nonparametric methodology for prediction problems. A standard way to use RFs includes generating a global RF to predict all test cases of interest. In this article, we propose growing different RFs specific to different test cases, namely case-specific random forests (CSRFs). In contrast to the bagging procedure in the building of standard RFs, the CSRF algorithm takes weighted bootstrap resamples to create individual trees, where we assign large weights to the training cases in close proximity to the test case of interest a priori. Tuning methods are discussed to avoid overfitting issues. Both simulation and real data examples show that the weighted bootstrap resampling used in CSRF construction can improve predictions for specific cases. We also propose a new case-specific variable importance (CSVI) measure as a way to compare the relative predictor variable importance for predicting a particular case. It is possible that the idea of building a predictor case-specifically can be generalized in other areas.