A two-stage gene selection scheme utilizing MRMR filter and GA wrapper

A two-stage gene selection scheme utilizing MRMR filter and GA wrapper
复制标题

DOI:
10.1007/s10115-010-0288-x
复制
发表时间:
2011-03-01
影响因子:
2.7
通讯作者:
Aboutajdine, Driss
Aboutajdine, Driss
中科院分区:
计算机科学4区
文献类型:
--
作者:
El Akadi, Ali;Amine, Aouatif;Aboutajdine, Driss

文献摘要

被引文献

相似文献

基因表达数据通常包含大量基因,但是少量样品。基因表达数据的特征选择旨在找到一组最能区分不同类型的生物样品的基因。在本文中,我们通过组合MRMR(最小冗余最大最大相关性)和GA(遗传算法)提出了一种两阶段选择算法。在第一阶段,MRMR用于过滤高维微阵列数据中的嘈杂和冗余基因。在第二阶段,GA使用分类器精度作为适应性函数来选择高度区分的基因。该方法在五个开放数据集上测试了肿瘤分类:NCI,淋巴瘤,肺,白血病和结肠使用支持载体机(SVM)和Na <Ve Bayes(NB)分类器。 MRMR-GA与MRMR滤波器和GA包装器的比较表明,我们的方法能够找到最小的基因子集,从而在剩下的交叉验证(LOOCV)方面具有最大的分类精度。
Gene expression data usually contain a large number of genes, but a small number of samples. Feature selection for gene expression data aims at finding a set of genes that best discriminates biological samples of different types. In this paper, we propose a two-stage selection algorithm for genomic data by combining MRMR (Minimum Redundancy-Maximum Relevance) and GA (Genetic Algorithm). In the first stage, MRMR is used to filter noisy and redundant genes in high-dimensional microarray data. In the second stage, the GA uses the classifier accuracy as a fitness function to select the highly discriminating genes. The proposed method is tested for tumor classification on five open datasets: NCI, Lymphoma, Lung, Leukemia and Colon using Support Vector Machine (SVM) and Na < ve Bayes (NB) classifiers. The comparison of the MRMR-GA with MRMR filter and GA wrapper shows that our method is able to find the smallest gene subset that gives the most classification accuracy in leave-one-out cross-validation (LOOCV).