Gene selection with guided regularized random forest

Gene selection with guided regularized random forest
复制标题

DOI:
10.1016/j.patcog.2013.05.018
复制
发表时间:
2013-12-01
影响因子:
8
通讯作者:
Runger, George
Runger, George
中科院分区:
计算机科学1区
文献类型:
--
作者:
Deng, Houtao;Runger, George

文献摘要

被引文献

相似文献

正则化随机森林(RRF)是最近提出的一种特征选择方法,它只建立一个集成。在RRF中,在每个树节点的训练数据的一部分上对特征进行评估。我们推导了一个节点中不同基尼信息增益值的个数的上界,并证明了多个特征可以在一个节点上以较少的实例和大量的特征共享相同的信息增益。因此,在实例数较少的节点中,RRF很可能选择相关性不强的特征,本文提出了一种增强型RRF,称为导引RRF(GRRF)。在GRRF中,来自普通随机森林(RF)的重要性得分被用来指导RRF的特征选择过程。在10个基因数据集上的实验表明,GRRF在参数变化时,总体上比RRF具有更好的精度性能。与RRF、varSelRF和Lasso Logistic回归(由RE分类器进行评估)相比,GRRF具有计算效率高、可以选择紧凑的特征子集并且具有与之相当的精度性能。此外,对于这里考虑的大多数数据集,应用于RRF选择的具有最小正则化的特征的RF优于应用于所有特征的RF。因此,如果精度被认为比特征子集的大小更重要,则可以考虑具有最小正则化的RRF。我们使用强分类器RF的精度性能来评估特征选择方法,并说明弱分类器捕获特征子集中包含的信息的能力较差。RRF和GRRF都是在“RRF”R包中实现的,可以在CRAN(官方R包存档)上找到。(C)2013爱思唯尔有限公司。保留所有权利。
The regularized random forest (RRF) was recently proposed for feature selection by building only one ensemble. In RRF the features are evaluated on a part of the training data at each tree node. We derive an upper bound for the number of distinct Gini information gain values in a node, and show that many features can share the same information gain at a node with a small number of instances and a large number of features. Therefore, in a node with a small number of instances, RRF is likely to select a feature not strongly relevant.Here an enhanced RRF, referred to as the guided RRF (GRRF), is proposed. In GRRF, the importance scores from an ordinary random forest (RF) are used to guide the feature selection process in RRF. Experiments on 10 gene data sets show that the accuracy performance of GRRF is, in general, more robust than RRF when their parameters change. GRRF is computationally efficient, can select compact feature subsets, and has competitive accuracy performance, compared to RRF, varSelRF and LASSO logistic regression (with evaluations from an RE classifier). Also, RF applied to the features selected by RRF with the minimal regularization outperforms RF applied to all the features for most of the data sets considered here. Therefore, if accuracy is considered more important than the size of the feature subset, RRF with the minimal regularization may be considered. We use the accuracy performance of RF, a strong classifier, to evaluate feature selection methods, and illustrate that weak classifiers are less capable of capturing the information contained in a feature subset. Both RRF and GRRF were implemented in the "RRF" R package available at CRAN, the official R package archive. (C) 2013 Elsevier Ltd. All rights reserved.