Bayesian mixture model based clustering of replicated microarray data

Bayesian mixture model based clustering of replicated microarray data
复制标题

DOI:
10.1093/bioinformatics/bth068
复制
发表时间:
2004-05-22
期刊:
影响因子:
5.8
通讯作者:
Bumgarner, RE
Bumgarner, RE
中科院分区:
生物学3区
文献类型:
--
作者:
Medvedovic, M;Yeung, KY;Bumgarner, RE

文献摘要

被引文献

相似文献

动机:通过聚类分析识别微阵列数据中的共表达模式已经成为揭示正在研究的生物过程的分子机制的一种有效方法。使用实验重复通常可以通过减少测量的实验变异性来提高聚类分析的精度。在这种情况下,贝叶斯混合物允许一个有效的利用信息,精确建模之间的重复variability.Results:我们开发了不同的变种贝叶斯混合物为基础的聚类程序与实验重复的基因表达数据的聚类。在这种方法中,微阵列数据的统计分布由贝叶斯混合模型描述。从聚类的后验分布创建共表达基因的聚类,该后验分布由Gibbs采样器估计。我们定义了具有不同重复间方差结构的无限和有限贝叶斯混合模型,并通过分析合成和真实世界的数据集来研究它们的效用。我们的分析结果表明:(1)当重复间变异性高时,通过仅进行两次实验重复获得的精度提高可能是显着的,(2)基因内变异性的精确建模对于准确识别共表达基因很重要,(3)具有“椭圆”重复间方差结构的无限混合模型总体上优于任何其他测试方法。我们还介绍了一个启发式修改的吉布斯采样器的基础上的“反向退火”的原则。这种修改有效地克服了吉布斯采样器从不同的初始位置开始时收敛到后验分布的不同模式的趋势。最后,我们证明了具有“椭圆”方差结构的贝叶斯无限混合模型能够在不知道“正确”聚类数的情况下识别数据的底层结构。
Motivation: Identifying patterns of co-expression in microarray data by cluster analysis has been a productive approach to uncovering molecular mechanisms underlying biological processes under investigation. Using experimental replicates can generally improve the precision of the cluster analysis by reducing the experimental variability of measurements. In such situations, Bayesian mixtures allow for an efficient use of information by precisely modeling between-replicates variability.Results: We developed different variants of Bayesian mixture based clustering procedures for clustering gene expression data with experimental replicates. In this approach, the statistical distribution of microarray data is described by a Bayesian mixture model. Clusters of co-expressed genes are created from the posterior distribution of clusterings, which is estimated by a Gibbs sampler. We define infinite and finite Bayesian mixture models with different between-replicates variance structures and investigate their utility by analyzing synthetic and the real-world datasets. Results of our analyses demonstrate that (1) improvements in precision achieved by performing only two experimental replicates can be dramatic when the between-replicates variability is high, (2) precise modeling of intra-gene variability is important for accurate identification of co-expressed genes and (3) the infinite mixture model with the 'elliptical' between-replicates variance structure performed overall better than any other method tested. We also introduce a heuristic modification to the Gibbs sampler based on the 'reverse annealing' principle. This modification effectively overcomes the tendency of the Gibbs sampler to converge to different modes of the posterior distribution when started from different initial positions. Finally, we demonstrate that the Bayesian infinite mixture model with 'elliptical' variance structure is capable of identifying the underlying structure of the data without knowing the 'correct' number of clusters.