Variable ranking and selection with random forest for unbalanced data

Variable ranking and selection with random forest for unbalanced data
复制标题

针对不平衡数据的随机森林变量排序和选择

DOI:
10.1017/eds.2022.34
复制
发表时间:
2022
期刊:
Environmental Data Science
影响因子:
--
通讯作者:
Bradter U
Bradter U
中科院分区:
--
文献类型:
--
作者:
Bradter U

文献摘要

参考文献

被引文献

相似文献

When one or several classes are much less prevalent than another class (unbalanced data), class error rates and variable importances of the machine learning algorithm random forest can be biased, particularly when sample sizes are smaller, imbalance levels higher, and effect sizes of important variables smaller. Using simulated data varying in size, imbalance level, number of true variables, their effect sizes, and the strength of multicollinearity between covariates, we evaluated how eight versions of random forest ranked and selected true variables out of a large number of covariates despite class imbalance. The version that calculated variable importance based on the area under the curve (AUC) was least adversely affected by class imbalance. For the same number of true variables, effect sizes, and multicollinearity between covariates, the AUC variable importance ranked true variables still highly at the lower sample sizes and higher imbalance levels at which the other seven versions no longer achieved high ranks for true variables. Conversely, using the Hellinger distance to split trees or downsampling the majority class already ranked true variables lower and more variably at the larger sample sizes and lower imbalance levels at which the other algorithms still ranked true variables highly. In variable selection, a higher proportion of true variables were identified when covariates were ranked by AUC importances and the proportion increased further when the AUC was used as the criterion in forward variable selection. In three case studies, known species–habitat relationships and their spatial scales were identified despite unbalanced data.
DOI: 10.1007/s10618-011-0222-1
发表时间: 2012-01-01
影响因子: 4.8
作者:
Cieslak, David A.;Hoens, T. Ryan;Kegelmeyer, W. Philip
通讯作者: Kegelmeyer, W. Philip
海林格距离作为平衡和不平衡分类数据集中随机森林分裂度量的研究
DOI: --
发表时间: 2020
影响因子: 8.5
作者:
R. Aler;J. Valls;Henrik Boström
通讯作者: Henrik Boström
使用随机森林和光学和雷达卫星数据绘制高地植被地图
DOI: 10.1002/rse2.32
发表时间: 2016
影响因子: 5.5
作者:
B. Barrett;Christoph Raab;F. Cawkwell;S. Green
通讯作者: S. Green
使用分层空间计数模型将当地物种-栖息地关系扩展到更大的景观
DOI: 10.1007/s10980-006-9005-2
发表时间: 2007
期刊: Landscape Ecology
影响因子: 5.2
作者:
W. Thogmartin;M. Knutson
通讯作者: M. Knutson
盐沼的轻度放牧增加了红脚鹬筑巢地点的可用性,但降低了它们的质量
DOI: --
发表时间: 2016
期刊:
影响因子: --
作者:
Elwyn Sharps;A. Garbutt;J. Hiddink;J. Smart;M. Skov
通讯作者: M. Skov