Enriched random forests

Enriched random forests
复制标题

DOI:
10.1093/bioinformatics/btn356
复制
发表时间:
2008-09-15
期刊:
影响因子:
5.8
通讯作者:
Lee, Yung-Seop
Lee, Yung-Seop
中科院分区:
生物学3区
文献类型:
--
作者:
Amaratunga, Dhammika;Cabrera, Javier;Lee, Yung-Seop

文献摘要

被引文献

相似文献

虽然随机森林分类过程在具有许多特征的数据集中工作得很好,但当特征数量巨大而真正提供信息的特征的百分比很小时,例如对于DNA微阵列数据,其性能往往会显著下降。在这种情况下,可以通过减少其节点由非信息性特征填充的树的贡献来改进该过程。在某种程度上,这可以通过预滤波来实现,但我们提出了一种新的、简单的、具有明显优越性能的调整:在每个节点上通过加权随机抽样而不是简单的随机抽样来选择合格的子集,并且权值向信息特征倾斜。这导致了丰富的随机森林。我们在几个实际的微阵列数据集中展示了该方法的优越性能。
Although the random forest classification procedure works well in datasets with many features, when the number of features is huge and the percentage of truly informative features is small, such as with DNA microarray data, its performance tends to decline significantly. In such instances, the procedure can be improved by reducing the contribution of trees whose nodes are populated by non-informative features. To some extent, this can be achieved by prefiltering, but we propose a novel, yet simple, adjustment that has demonstrably superior performance: choose the eligible subsets at each node by weighted random sampling instead of simple random sampling, with the weights tilted in favor of the informative features. This results in an enriched random forest. We illustrate the superior performance of this procedure in several actual microarray datasets.