Consensus clustering: A resampling-based method for class discovery and visualization of gene expression microarray data

Consensus clustering: A resampling-based method for class discovery and visualization of gene expression microarray data
复制标题

DOI:
10.1023/a:1023949509487
复制
发表时间:
2003-07-01
期刊:
影响因子:
7.5
通讯作者:
Golub, T
Golub, T
中科院分区:
计算机科学3区
文献类型:
--
作者:
Monti, S;Tamayo, P;Golub, T

文献摘要

被引文献

相似文献

在本文中,我们提出了一种新的方法,类发现和聚类验证的任务,分析基因表达数据。该方法最好被认为是一种分析方法,以指导和帮助使用各种可用的聚类算法。我们称之为新的方法共识聚类,并结合resception技术,它提供了一种方法来表示跨多个运行的聚类算法的共识,并评估所发现的集群的稳定性。该方法还可以用于表示具有随机重启的聚类算法(例如K-means、基于模型的贝叶斯聚类、SOM等)的多次运行的一致性,以考虑其对初始条件的敏感性。最后,它提供了一个可视化工具来检查集群数量、成员关系和边界。我们提出了我们的模拟数据和真实的基因表达数据的实验结果,旨在评估该方法在发现生物学上有意义的集群的有效性。
In this paper we present a new methodology of class discovery and clustering validation tailored to the task of analyzing gene expression data. The method can best be thought of as an analysis approach, to guide and assist in the use of any of a wide range of available clustering algorithms. We call the new methodology consensus clustering, and in conjunction with resampling techniques, it provides for a method to represent the consensus across multiple runs of a clustering algorithm and to assess the stability of the discovered clusters. The method can also be used to represent the consensus over multiple runs of a clustering algorithm with random restart ( such as K-means, model-based Bayesian clustering, SOM, etc.), so as to account for its sensitivity to the initial conditions. Finally, it provides for a visualization tool to inspect cluster number, membership, and boundaries. We present the results of our experiments on both simulated data and real gene expression data aimed at evaluating the effectiveness of the methodology in discovering biologically meaningful clusters.