GeneMCL in microarray analysis

GeneMCL in microarray analysis
复制标题

DOI:
10.1016/j.compbiolchem.2005.07.002
复制
发表时间:
2005-10-01
影响因子:
3.1
通讯作者:
Crabbe, MJC
Crabbe, MJC
中科院分区:
生物学3区
文献类型:
--
作者:
Lattimore, BS;van Dongen, S;Crabbe, MJC

文献摘要

被引文献

相似文献

准确,可靠地识别出具有基因表达概况数据集的实际簇数,当没有关于群集结构的其他信息时,很少有算法解决的问题。 GeneMCL将微阵列分析数据转换为由边缘连接的节点组成的图,其中节点代表基因,边缘代表这些基因表达的相似性,如接近度测量所给出。该测量被认为是Pearson相关系数与局部非线性重新缩放步骤相结合。所得图是输入到Markov群集(MCL)算法的,该算法是一种优雅,确定性,非特异性和可扩展的方法,该方法模拟了整个图的随机流。该算法固有地受到存在的任何群集结构的影响,并将图形迅速分解为凝聚群。 GenEMCL算法的潜力用van't Veer乳腺癌数据库的5730基因子集(IG)证明,该数据库显示出来反映了基本的生物学机制。 (c)2005 Elsevier Ltd.保留所有权利。
Accurately and reliably identifying the actual number of clusters present with a dataset of gene expression profiles, when no additional information on cluster structure is available, is a problem addressed by few algorithms. GeneMCL transforms microarray analysis data into a graph consisting of nodes connected by edges, where the nodes represent genes, and the edges represent the similarity in expression of those genes, as given by a proximity measurement. This measurement is taken to be the Pearson correlation coefficient combined with a local non-linear rescaling step. The resulting graph is input to the Markov Cluster (MCL) algorithm, which is an elegant, deterministic, non-specific and scalable method, which models stochastic flow through the graph. The algorithm is inherently affected by any cluster structure present, and rapidly decomposes a graph into cohesive clusters. The potential of the GeneMCL algorithm is demonstrated with a 5730 gene subset (IGS) of the Van't Veer breast cancer database, for which the clusterings are shown to reflect underlying biological mechanisms. (c) 2005 Elsevier Ltd. All rights reserved.