Unsupervised feature selection via two-way ordering in gene expression analysis

Unsupervised feature selection via two-way ordering in gene expression analysis
复制标题

DOI:
10.1093/bioinformatics/btg149
复制
发表时间:
2003-07-01
期刊:
影响因子:
5.8
通讯作者:
Ding, CHQ
Ding, CHQ
中科院分区:
生物学3区
文献类型:
--
作者:
Ding, CHQ

文献摘要

被引文献

相似文献

动机:选择与某些表型最相关和信息最丰富的基因是基因表达分析的一个重要方面。目前大多数方法选择基因的基础上已知的表型信息。然而,某些组的基因可能对应于新的表型,这是未知的,它是重要的,以开发新的有效的选择方法,他们的发现,而不使用任何先前的表型informations.Results:我们提出并研究了一种新的方法来选择相关基因的基础上,他们的相似性信息。该方法依赖于丢弃不相关基因的机制。基因表达数据的双向排序可以迫使不相关的基因朝向排序中的中间并且因此可以被丢弃。基于方差和主成分分析的机制也进行了研究。当应用于结肠癌和白血病的表达谱时,无监督方法优于简单使用所有基因的基线算法,并且它还选择了与使用监督方法选择的基因接近的相关基因。
Motivation: Selection of genes most relevant and informative for certain phenotypes is an important aspect in gene expression analysis. Most current methods select genes based on known phenotype information. However, certain set of genes may correspond to new phenotypes which are yet unknown, and it is important to develop novel effective selection methods for their discovery without using any prior phenotype information.Results: We propose and study a new method to select relevant genes based on their similarity information only. The method relies on a mechanism for discarding irrelevant genes. A two-way ordering of gene expression data can force irrelevant genes towards the middle in the ordering and thus can be discarded. Mechanisms based on variance and principal component analysis are also studied. When applied to expression profiles of colon cancer and leukemia, the unsupervised method outperforms the baseline algorithm that simply uses all genes, and it also selects relevant genes close to those selected using supervised methods.