Robust multi-scale clustering of large DNA microarray datasets with the consensus algorithm

Robust multi-scale clustering of large DNA microarray datasets with the consensus algorithm
复制标题

DOI:
10.1093/bioinformatics/bti746
复制
发表时间:
2006-01-01
期刊:
影响因子:
5.8
通讯作者:
Hansen, LK
Hansen, LK
中科院分区:
生物学3区
文献类型:
--
作者:
Grotkjær, T;Winther, O;Hansen, LK

文献摘要

被引文献

相似文献

动机:层次聚类和重定位聚类(例如 K 均值和自组织图)已成为显示和分析全基因组 DNA 微阵列表达数据的成功工具。然而,层次聚类的结果对异常值很敏感,并且大多数重定位方法给出的结果取决于算法的初始化。因此,很难评估结果的意义。我们开发了一种共识聚类算法,其中最终结果是多次聚类运行的平均值,从而提供稳健且可重复的聚类,能够捕获小信号变化。该算法保留了层次聚类的宝贵特性,这对于结果的可视化和解释非常有用。结果:我们首次表明,通过收集共现矩阵中的重复出现的聚类模式,可以在 DNA 微阵列分析中利用多次聚类运行。结果表明,使用高斯变分贝叶斯混合或 K 均值进行多次聚类获得的一致聚类显着降低了模拟数据集的分类错误率。该方法灵活,可以从不同的聚类算法中找到共识簇。因此,该算法可以作为定量测试不同聚类算法同质性的框架。我们将该方法与许多最先进的聚类方法进行比较。结果表明,该方法非常稳健,并且对于真实的模拟数据集分类错误率较低。该算法还针对真实数据集进行了演示。结果表明,无需保守统计或倍数变化排除数据,即可发现更具生物学意义的转录模式。
Motivation: Hierarchical and relocation clustering (e.g. K-means and self-organizing maps) have been successful tools in the display and analysis of whole genome DNA microarray expression data. However, the results of hierarchical clustering are sensitive to outliers, and most relocation methods give results which are dependent on the initialization of the algorithm. Therefore, it is difficult to assess the significance of the results. We have developed a consensus clustering algorithm, where the final result is averaged over multiple clustering runs, giving a robust and reproducible clustering, capable of capturing small signal variations. The algorithm preserves valuable properties of hierarchical clustering, which is useful for visualization and interpretation of the results.Results: We show for the first time that one can take advantage of multiple clustering runs in DNA microarray analysis by collecting re-occurring clustering patterns in a co-occurrence matrix. The results show that consensus clustering obtained from clustering multiple times with Variational Bayes Mixtures of Gaussians or K-means significantly reduces the classification error rate for a simulated dataset. The method is flexible and it is possible to find consensus clusters from different clustering algorithms. Thus, the algorithm can be used as a framework to test in a quantitative manner the homogeneity of different clustering algorithms. We compare the method with a number of state-of-the-art clustering methods. It is shown that the method is robust and gives low classification error rates for a realistic, simulated dataset. The algorithm is also demonstrated for real datasets. It is shown that more biological meaningful transcriptional patterns can be found without conservative statistical or fold-change exclusion of data.