Enriched Random Forest for High Dimensional Genomic Data.

Enriched Random Forest for High Dimensional Genomic Data.
复制标题

高维基因组数据的富集随机森林。

DOI:
10.1109/tcbb.2021.3089417
复制
发表时间:
2022-09
期刊:
IEEE/ACM transactions on computational biology and bioinformatics
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

被引文献

相似文献

随机森林等集成方法在高维数据集上效果很好。然而,当特征数量与样本数量相比非常大且真正提供信息的特征的百分比非常小时,传统随机森林的性能显着下降。为此,我们开发了一种新颖的方法,通过减少节点填充信息量较少的特征的树的贡献来增强传统随机森林的性能。所提出的方法通过加权随机采样在每个节点选择合格的子集,而不是传统随机森林中的简单随机采样。我们将这种改进的随机森林算法称为“丰富随机森林”。使用几个高维微阵列数据集,我们评估了我们的方法在回归和分类设置中的性能。此外,我们还证明了平衡留一交叉验证在计算特征权重时减少计算负载和减少样本大小的有效性。总体而言,结果表明,丰富的随机森林提高了传统随机森林的预测精度,特别是当相关特征很少时。
Ensemble methods such as random forest works well on high-dimensional datasets. However, when the number of features is extremely large compared to the number of samples and the percentage of truly informative feature is very small, performance of traditional random forest decline significantly. To this end, we develop a novel approach that enhance the performance of traditional random forest by reducing the contribution of trees whose nodes are populated with less informative features. The proposed method selects eligible subsets at each node by weighted random sampling as opposed to simple random sampling in traditional random forest. We refer to this modified random forest algorithm as ”Enriched Random Forest”. Using several high-dimensional micro-array datasets, we evaluate the performance of our approach in both regression and classification settings. In addition, we also demonstrate the effectiveness of balanced leave-one-out cross-validation to reduce computational load and decrease sample size while computing feature weights. Overall, the results indicate that enriched random forest improves the prediction accuracy of traditional random forest, especially when relevant features are very few.