Bayesian infinite mixture model based clustering of gene expression profiles

Bayesian infinite mixture model based clustering of gene expression profiles
复制标题

DOI:
10.1093/bioinformatics/18.9.1194
复制
发表时间:
2002-09-01
期刊:
影响因子:
5.8
通讯作者:
Sivaganesan, S
Sivaganesan, S
中科院分区:
生物学3区
文献类型:
--
作者:
Medvedovic, M;Sivaganesan, S

文献摘要

被引文献

相似文献

动机:通过对微阵列实验中产生的基因表达数据进行聚类分析所获得的结果的生物学意义已经在许多研究中得到证实。在这篇文章中,我们专注于发展的聚类过程的基础上的概念贝叶斯模型平均和精确的统计模型的expression data.Results:我们开发了一个聚类过程的基础上贝叶斯无限混合模型,并将其应用到聚类基因表达profiles。具有相似表达模式的基因簇从由随机数据生成模型隐式定义的聚类的后验分布中识别。聚类的后验分布由Gibbs抽样器估计。我们通过计算共表达的后验成对概率来总结聚类的后验分布,并使用完全连锁原理来创建聚类。这种方法比通常的聚类程序有几个优点。该分析允许纳入一个合理的概率模型生成数据。该方法不需要指定的集群的数量和得到的最佳聚类是通过平均模型与所有可能的集群数量。自动检测与任何其他谱不相似的表达谱,该方法结合了实验重复,并且可以扩展以容纳缺失的数据。这种方法代表了表达数据的基于模型的聚类分析的质的转变,因为它允许在表达谱相似性的置信度的最终评估中纳入模型选择中涉及的不确定性。我们还证明了将实验变异性的信息纳入聚类模型的重要性。
Motivation: The biologic significance of results obtained through cluster analyses of gene expression data generated in microarray experiments have been demonstrated in many studies. In this article we focus on the development of a clustering procedure based on the concept of Bayesian model-averaging and a precise statistical model of expression data.Results: We developed a clustering procedure based on the Bayesian infinite mixture model and applied it to clustering gene expression profiles. Clusters of genes with similar expression patterns are identified from the posterior distribution of clusterings defined implicitly by the stochastic data-generation model. The posterior distribution of clusterings is estimated by a Gibbs sampler. We summarized the posterior distribution of clusterings by calculating posterior pairwise probabilities of co-expression and used the complete linkage principle to create clusters. This approach has several advantages over usual clustering procedures. The analysis allows for incorporation of a reasonable probabilistic model for generating data. The method does not require specifying the number of clusters and resulting optimal clustering is obtained by averaging over models with all possible numbers of clusters. Expression profiles that are not similar to any other profile are automatically detected, the method incorporates experimental replicates, and it can be extended to accommodate missing data. This approach represents a qualitative shift in the model-based cluster analysis of expression data because it allows for incorporation of uncertainties involved in the model selection in the final assessment of confidence in similarities of expression profiles. We also demonstrated the importance of incorporating the information on experimental variability into the clustering model.