BAYESIAN CLUSTERING OF REPLICATED TIME-COURSE GENE EXPRESSION DATA WITH WEAK SIGNALS

BAYESIAN CLUSTERING OF REPLICATED TIME-COURSE GENE EXPRESSION DATA WITH WEAK SIGNALS
复制标题

DOI:
10.1214/13-aoas650
复制
发表时间:
2013-09-01
影响因子:
1.8
通讯作者:
Tavare, Simon
Tavare, Simon
中科院分区:
数学4区
文献类型:
--
作者:
Fu, Audrey Qiuyan;Russell, Steven;Tavare, Simon

文献摘要

被引文献

相似文献

为了识别新的基因表达的动态模式,我们开发了一种统计方法来聚类噪声测量的基因表达收集从多个重复在多个时间点,与一个未知数量的集群。我们提出了一个随机效应的混合模型,再加上一个Dirichlet过程前聚类。混合模型公式化允许概率聚类分配。随机效应公式允许将数据的总变异性归因于与实验设计一致的来源,特别是当噪声水平很高且时间依赖性不强时。Dirichlet过程先验在分区上引入先验分布,并有助于从数据中估计聚类(或混合成分)的数量。我们进一步解决了与Dirichlet过程先验方法相关的两个挑战。一是高效采样。我们开发了一种新的大都会黑斯廷斯马尔可夫链蒙特卡罗(MCMC)程序来采样的分区。另一个是有效地使用MCMC样本形成集群。我们提出了一个两步的后验推理过程,其中包括重新标记和重新标记,估计后验分配概率矩阵。该矩阵可以直接用于聚类分配,同时描述聚类中的不确定性。我们通过模拟数据证明了我们的模型和抽样程序的有效性。将我们的方法应用于从果蝇成体肌肉细胞中收集的真实的数据集,经过5分钟的Notch激活,我们在163个差异表达基因中识别出14个不同的转录反应簇,这为Notch信号通路中潜在的转录机制提供了新的见解。这里开发的算法在R包DIRECT中实现,可在CRAN上获得。
To identify novel dynamic patterns of gene expression, we develop a statistical method to cluster noisy measurements of gene expression collected from multiple replicates at multiple time points, with an unknown number of clusters. We propose a random-effects mixture model coupled with a Dirichlet-process prior for clustering. The mixture model formulation allows for probabilistic cluster assignments. The random-effects formulation allows for attributing the total variability in the data to the sources that are consistent with the experimental design, particularly when the noise level is high and the temporal dependence is not strong. The Dirichlet-process prior induces a prior distribution on partitions and helps to estimate the number of clusters (or mixture components) from the data. We further tackle two challenges associated with Dirichlet-process prior-based methods. One is efficient sampling. We develop a novel Metropolis-Hastings Markov Chain Monte Carlo (MCMC) procedure to sample the partitions. The other is efficient use of the MCMC samples in forming clusters. We propose a two-step procedure for posterior inference, which involves resampling and relabeling, to estimate the posterior allocation probability matrix. This matrix can be directly used in cluster assignments, while describing the uncertainty in clustering. We demonstrate the effectiveness of our model and sampling procedure through simulated data. Applying our method to a real data set collected from Drosophila adult muscle cells after five-minute Notch activation, we identify 14 clusters of different transcriptional responses among 163 differentially expressed genes, which provides novel insights into underlying transcriptional mechanisms in the Notch signaling pathway. The algorithm developed here is implemented in the R package DIRECT, available on CRAN.