HykGene: a hybrid approach for selecting marker genes for phenotype classification using microarray gene expression data

HykGene: a hybrid approach for selecting marker genes for phenotype classification using microarray gene expression data
复制标题

DOI:
10.1093/bioinformatics/bti192
复制
发表时间:
2005-04-15
期刊:
影响因子:
5.8
通讯作者:
Pearlman, J
Pearlman, J
中科院分区:
生物学3区
文献类型:
--
作者:
Wang, YH;Makedon, FS;Pearlman, J

文献摘要

被引文献

相似文献

动机:最近的研究表明,微阵列基因表达数据是有用的许多疾病的表型分类。这种分类的一个主要问题是特征(基因)的数量大大超过了实例(组织样本)的数量。它已被证明,选择一个小集合的信息基因可以导致提高分类精度。已经提出了许多方法来解决这个基因选择问题。大多数先前的基因排序方法通常选择50-200个排名靠前的基因,并且这些基因通常高度相关。我们的目标是选择一个小的非冗余标记基因,是最相关的classification task.Results:为了实现这一目标,我们开发了一种新的混合方法,结合基因排序和聚类分析。在这种方法中,我们首先应用特征过滤算法来选择一组排名靠前的基因,然后对这些基因应用层次聚类来生成树状图。最后,通过扫描线算法对树状图进行分析,并通过折叠密集簇来选择标记基因。使用三个公共数据集的实证研究表明,我们的方法是能够选择相对较少的标记基因,同时提供相同或更好的留一交叉验证的准确性相比,直接使用排名靠前的基因进行分类的方法。
Motivation: Recent studies have shown that microarray gene expression data are useful for phenotype classification of many diseases. A major problem in this classification is that the number of features (genes) greatly exceeds the number of instances (tissue samples). It has been shown that selecting a small set of informative genes can lead to improved classification accuracy. Many approaches have been proposed for this gene selection problem. Most of the previous gene ranking methods typically select 50-200 top-ranked genes and these genes are often highly correlated. Our goal is to select a small set of non-redundant marker genes that are most relevant for the classification task.Results: To achieve this goal, we developed a novel hybrid approach that combines gene ranking and clustering analysis. In this approach, we first applied feature filtering algorithms to select a set of top-ranked genes, and then applied hierarchical clustering on these genes to generate a dendrogram. Finally, the dendrogram was analyzed by a sweep-line algorithm and marker genes are selected by collapsing dense clusters. Empirical study using three public datasets shows that our approach is capable of selecting relatively few marker genes while offering the same or better leave-one-out cross-validation accuracy compared with approaches that use top-ranked genes directly for classification.