A Top-r Feature Selection Algorithm for Microarray Gene Expression Data

A Top-r Feature Selection Algorithm for Microarray Gene Expression Data
复制标题

DOI:
10.1109/tcbb.2011.151
复制
发表时间:
2012-05-01
影响因子:
4.5
通讯作者:
Miyano, Satoru
Miyano, Satoru
中科院分区:
工程技术3区
文献类型:
--
作者:
Sharma, Alok;Imoto, Seiya;Miyano, Satoru

文献摘要

被引文献

相似文献

大多数传统的特征选择算法都有一个缺点,即在适当的基因子集中可以在分类精度方面表现良好的弱排序基因将被排除在选择之外。针对这一不足,我们提出了一种基因表达数据分析样本分类中的特征选择算法。该算法首先将基因划分为大小相对较小的子集(大小约为h),然后从一个子集中选择信息量较小的基因子集(大小为r < h),并将所选基因与另一个基因子集(大小为r)合并以更新基因子集。我们重复这个过程,直到所有子集合并成一个信息子集。我们通过分析三个不同的基因表达数据集来说明所提出算法的有效性。我们的方法对所有测试数据集都显示出很好的分类精度。我们还展示了所选基因在生物学功能方面的相关性。
Most of the conventional feature selection algorithms have a drawback whereby a weakly ranked gene that could perform well in terms of classification accuracy with an appropriate subset of genes will be left out of the selection. Considering this shortcoming, we propose a feature selection algorithm in gene expression data analysis of sample classifications. The proposed algorithm first divides genes into subsets, the sizes of which are relatively small (roughly of size h), then selects informative smaller subsets of genes (of size r < h) from a subset and merges the chosen genes with another gene subset (of size r) to update the gene subset. We repeat this process until all subsets are merged into one informative subset. We illustrate the effectiveness of the proposed algorithm by analyzing three distinct gene expression data sets. Our method shows promising classification accuracy for all the test data sets. We also show the relevance of the selected genes in terms of their biological functions.