Bayesian Inference for Gene Expression and Proteomics: Model-Based Clustering for Expression Data via a Dirichlet Process Mixture Model

Bayesian Inference for Gene Expression and Proteomics: Model-Based Clustering for Expression Data via a Dirichlet Process Mixture Model
复制标题

基因表达和蛋白质组学的贝叶斯推理:通过狄利克雷过程混合模型对表达数据进行基于模型的聚类

DOI:
--
复制
发表时间:
2006
期刊:
影响因子:
--
通讯作者:
D. B. Dahl
D. B. Dahl
中科院分区:
--
文献类型:
--
作者:
D. B. Dahl

文献摘要

参考文献

被引文献

相似文献

本章描述了基于明确定义的统计模型的微阵列表达数据的聚类过程,具体地说,共轭Dirichlet过程混合模型。聚类算法将控制表达的潜在变量相等的基因分组,即属于相同混合成分的基因。该模型采用马尔可夫链蒙特卡罗方法进行拟合,并利用共轭特性减轻了计算负担。本章介绍了一种基于最小二乘距离从两个基因被聚类的后验概率得到真实聚类的点估计的方法。与特定的聚类方法不同,该模型提供了关于聚类的不确定性度量。此外,该模型自动估计聚类的数量,并量化关于这一重要参数的不确定性。在仿真研究中,将该方法与其他聚类方法进行了比较。最后,用实际的微阵列数据对该方法进行了验证。
This chapter describes a clustering procedure for microarray expression data based on a well-defined statistical model, specifically, a conjugate Dirichlet process mixture model. The clustering algorithm groups genes whose latent variables governing expression are equal, that is, genes belonging to the same mixture component. The model is fit with Markov chain Monte Carlo and the computational burden is eased by exploiting conjugacy. This chapter introduces a method to get a point estimate of the true clustering based on least-squares distances from the posterior probability that two genes are clustered. Unlike ad hoc clustering methods, the model provides measures of uncertainty about the clustering. Further, the model automatically estimates the number of clusters and quantifies uncertainty about this important parameter. The method is compared to other clustering methods in a simulation study. Finally, the method is demonstrated with actual microarray data.
DOI: 10.1152/physiolgenomics.00172.2002
发表时间: 2003-04-16
影响因子: 4.6
作者:
Edwards, MG;Sarkar, D;Prolla, TA
通讯作者: Prolla, TA