Bayesian models for pooling microarray studies with multiple sources of replications

Bayesian models for pooling microarray studies with multiple sources of replications
复制标题

DOI:
10.1186/1471-2105-7-247
复制
发表时间:
2006-05-05
期刊:
影响因子:
3
通讯作者:
Liu, Jun S.
Liu, Jun S.
中科院分区:
生物学4区
文献类型:
--
作者:
Conlon, Erin M.;Song, Joon J.;Liu, Jun S.

文献摘要

被引文献

相似文献

背景:生物学家经常进行多个但不同的cDNA微阵列研究,这些研究都针对相同的生物系统或途径。在每项研究中,通常会在重复的相同实验中制作重复切片。跨研究汇集信息可以帮助更准确地识别真正的靶基因。在这里,我们介绍了一种方法来整合多个独立的studies.Results:我们引入贝叶斯层次模型池跨多个独立的研究,以确定高表达的基因的cDNA微阵列数据。每项研究都有多个变异来源,即在重复的相同实验中重复切片。我们的模型产生了基因特异性差异表达的后验概率,这提供了一种直接的方法来对基因进行排序,并提供了错误发现率(FDR)的贝叶斯估计。在模拟结合两个和五个独立的研究,与固定的FDR水平,我们观察到发现的基因的数量大幅增加,在合并与个别分析。当输出基因的数量固定时(e.例如,在一个实施例中,前100名),合并模型发现明显更多的真正差异表达的基因比个别研究。我们还能够确定更多的差异表达的基因,从合并两个独立的研究在枯草芽孢杆菌比从每个单独的数据集。最后,我们观察到,在我们的模拟研究中,我们的贝叶斯FDR估计跟踪真实的FDR very well.Conclusion:我们的方法提供了一个有凝聚力的框架,结合多个,但不相同的微阵列研究与几个来源的复制,从同一平台产生的数据。我们假设每个研究只包含两个条件:实验样本和对照样本。我们证明了我们的模型适用于少量的研究,这些研究要么是预缩放的,要么没有离群值。
Background: Biologists often conduct multiple but different cDNA microarray studies that all target the same biological system or pathway. Within each study, replicate slides within repeated identical experiments are often produced. Pooling information across studies can help more accurately identify true target genes. Here, we introduce a method to integrate multiple independent studies efficiently.Results: We introduce a Bayesian hierarchical model to pool cDNA microarray data across multiple independent studies to identify highly expressed genes. Each study has multiple sources of variation, i.e. replicate slides within repeated identical experiments. Our model produces the genespecific posterior probability of differential expression, which provides a direct method for ranking genes, and provides Bayesian estimates of false discovery rates (FDR). In simulations combining two and five independent studies, with fixed FDR levels, we observed large increases in the number of discovered genes in pooled versus individual analyses. When the number of output genes is fixed ( e. g., top 100), the pooled model found appreciably more truly differentially expressed genes than the individual studies. We were also able to identify more differentially expressed genes from pooling two independent studies in Bacillus subtilis than from each individual data set. Finally, we observed that in our simulation studies our Bayesian FDR estimates tracked the true FDRs very well.Conclusion: Our method provides a cohesive framework for combining multiple but not identical microarray studies with several sources of replication, with data produced from the same platform. We assume that each study contains only two conditions: an experimental and a control sample. We demonstrated our model's suitability for a small number of studies that have been either pre-scaled or have no outliers.