Bayesian Inference for Gene Expression and Proteomics: Model-Based Clustering for Expression Data via a Dirichlet Process Mixture Model
Bayesian Inference for Gene Expression and Proteomics: Model-Based Clustering for Expression Data via a Dirichlet Process Mixture Model
复制标题
基因表达和蛋白质组学的贝叶斯推理:通过狄利克雷过程混合模型对表达数据进行基于模型的聚类
DOI:
--
复制
发表时间:
2006
期刊:
影响因子:
--
通讯作者:
D. B. Dahl
中科院分区:
文献类型:
--
作者:
D. B. Dahl
This chapter describes a clustering procedure for microarray expression data based on a well-defined statistical model, specifically, a conjugate Dirichlet process mixture model. The clustering algorithm groups genes whose latent variables governing expression are equal, that is, genes belonging to the same mixture component. The model is fit with Markov chain Monte Carlo and the computational burden is eased by exploiting conjugacy. This chapter introduces a method to get a point estimate of the true clustering based on least-squares distances from the posterior probability that two genes are clustered. Unlike ad hoc clustering methods, the model provides measures of uncertainty about the clustering. Further, the model automatically estimates the number of clusters and quantifies uncertainty about this important parameter. The method is compared to other clustering methods in a simulation study. Finally, the method is demonstrated with actual microarray data.
影响因子:
4.6
作者:
Edwards, MG;Sarkar, D;Prolla, TA
通讯作者:
Prolla, TA