Entropy-based gene ranking without selection bias for the predictive classification of microarray data.

Entropy-based gene ranking without selection bias for the predictive classification of microarray data.
复制标题

DOI:
10.1186/1471-2105-4-54
复制
发表时间:
2003-11-06
期刊:
影响因子:
3
通讯作者:
Jurman G
Jurman G
中科院分区:
生物学4区
文献类型:
--
作者:
Furlanello C;Serafini M;Merler S;Jurman G

文献摘要

参考文献

被引文献

相似文献

我们描述了E-RFE方法的基因排序,这是有用的阵列数据的预测分类中的标记的识别。该方法支持一个实用的建模方案,旨在避免构建分类规则的基础上选择太小的基因子集(被称为选择偏差的影响,其中估计的预测误差过于乐观,由于测试样本已经考虑在特征选择过程)。使用E-RFE,我们通过使用SVM权重分布的熵度量来消除不感兴趣的基因块,从而加快了SVM分类器的递归特征消除(RFE)。根据两层模型评估程序选择最佳基因子集:通过外部分层分区重采样方案复制建模,并且在每次运行中,使用内部K折交叉验证进行E-RFE排名。此外,最佳的基因数目可以估计根据Zipf定律轮廓的饱和度。在不降低分类精度的情况下,E-RFE相对于标准RFE允许100的加速因子,同时改进替代参数RFE减少策略。因此,基因选择和误差估计的过程是切实可行的,确保选择偏差的控制,并提供基因重要性的额外诊断指标。
We describe the E-RFE method for gene ranking, which is useful for the identification of markers in the predictive classification of array data. The method supports a practical modeling scheme designed to avoid the construction of classification rules based on the selection of too small gene subsets (an effect known as the selection bias, in which the estimated predictive errors are too optimistic due to testing on samples already considered in the feature selection process). With E-RFE, we speed up the recursive feature elimination (RFE) with SVM classifiers by eliminating chunks of uninteresting genes using an entropy measure of the SVM weights distribution. An optimal subset of genes is selected according to a two-strata model evaluation procedure: modeling is replicated by an external stratified-partition resampling scheme, and, within each run, an internal K-fold cross-validation is used for E-RFE ranking. Also, the optimal number of genes can be estimated according to the saturation of Zipf's law profiles. Without a decrease of classification accuracy, E-RFE allows a speed-up factor of 100 with respect to standard RFE, while improving on alternative parametric RFE reduction strategies. Thus, a process for gene selection and error estimation is made practical, ensuring control of the selection bias, and providing additional diagnostic indicators of gene importance.
DOI: 10.1093/bioinformatics/18.1.39
发表时间: 2002-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Nguyen, DV;Rocke, DM
通讯作者: Rocke, DM
DOI: 10.1103/physrevlett.90.088102
发表时间: 2003-02-28
影响因子: 8.6
作者:
Furusawa, C;Kaneko, K
通讯作者: Kaneko, K
DOI: 10.1038/35000501
发表时间: 2000-02-03
期刊: NATURE
影响因子: 64.8
作者:
Alizadeh, AA;Eisen, MB;Staudt, LM
通讯作者: Staudt, LM
DOI: 10.1006/jtbi.2002.3145
发表时间: 2002-12-21
影响因子: 2
作者:
Li, WT;Yang, YN
通讯作者: Yang, YN
DOI: 10.1084/jem.20021726
发表时间: 2003-06-02
期刊: The Journal of experimental medicine
影响因子: --
作者:
Kari L;Loboda A;Nebozhyn M;Rook AH;Vonderheid EC;Nichols C;Virok D;Chang C;Horng WH;Johnston J;Wysocka M;Showe MK;Showe LC
通讯作者: Showe LC