Variable importance-weighted Random Forests

Variable importance-weighted Random Forests
复制标题

DOI:
10.1007/s40484-017-0121-6
复制
发表时间:
2017-12-01
影响因子:
3.1
通讯作者:
Zhao, Hongyu
Zhao, Hongyu
中科院分区:
生物学4区
文献类型:
--
作者:
Liu, Yiyi;Zhao, Hongyu

文献摘要

被引文献

相似文献

背景:随机森林是一种流行的分类和回归方法,已被证明对于生物学研究中的各种预测问题非常有效。然而,当特征数量增加时,其性能往往会下降。为了解决这个限制,提出了特征消除随机森林,仅使用具有最大变量重要性得分的特征。然而,该方法的性能并不令人满意,可能是由于其严格的特征选择以及森林树木之间的相关性增加。方法:我们提出了变量重要性加权随机森林,它不是在每个节点以相等的概率采样特征来构建树,而是根据变量重要性分数对特征进行采样,然后从随机选择的特征中选择最佳分割。结果:我们通过全面的模拟和真实数据分析来评估我们的方法的性能,包括回归和分类。与标准随机森林和特征消除随机森林方法相比,我们提出的方法在大多数情况下都提高了性能。结论:通过将变量重要性得分纳入随机特征选择步骤,我们的方法可以更好地利用信息较多的特征,而不会完全忽略信息较少的特征,因此在存在弱信号和大噪声的情况下提高了预测精度。我们在原始R包“randomForest”的基础上实现了一个R包“viRandomForests”,它可以从http://zhaocenter.org/software免费下载。
Background: Random Forests is a popular classification and regression method that has proven powerful for various prediction problems in biological studies. However, its performance often deteriorates when the number of features increases. To address this limitation, feature elimination Random Forests was proposed that only uses features with the largest variable importance scores. Yet the performance of this method is not satisfying, possibly due to its rigid feature selection, and increased correlations between trees of forest. Methods: We propose variable importance-weighted Random Forests, which instead of sampling features with equal probability at each node to build up trees, samples features according to their variable importance scores, and then select the best split from the randomly selected features. Results: We evaluate the performance of our method through comprehensive simulation and real data analyses, for both regression and classification. Compared to the standard Random Forests and the feature elimination Random Forests methods, our proposed method has improved performance in most cases. Conclusions: By incorporating the variable importance scores into the random feature selection step, our method can better utilize more informative features without completely ignoring less informative ones, hence has improved prediction accuracy in the presence of weak signals and large noises. We have implemented an R package "viRandomForests" based on the original R package "randomForest" and it can be freely downloaded from http:// zhaocenter.org/software.