A Bayesian model for cross-study differential gene expression.

A Bayesian model for cross-study differential gene expression.
复制标题

DOI:
10.1198/jasa.2009.ap07611
复制
发表时间:
2009
影响因子:
3.7
通讯作者:
Nobel AB
Nobel AB
中科院分区:
数学1区
文献类型:
--
作者:
Scharpf RB;Tjelmeland H;Parmigiani G;Nobel AB

文献摘要

参考文献

被引文献

相似文献

在本文中,我们为从几个研究中收集的微阵列表达数据定义了一个分层贝叶斯模型,并使用它来识别在两个条件下显示差异表达的基因。主要特征包括跨基因和研究的收缩,以及允许平台之间的交互和估计效果的灵活建模,以及跨研究的一致和不一致的差异表达。我们使用人工数据和“分裂研究”验证方法,对模型的性能进行了全面的评估,这种方法不仅在零假设下,而且在现实的替代方案下,提供对模型行为的不可知性评估。人工数据的仿真结果表明了贝叶斯模型的优越性。贝叶斯模型的1-AuC值大约是t-统计量和SAM-统计量的直接组合的对应值的一半。此外,模拟还为贝叶斯模型何时最有用提供了指导。最值得注意的是,在小型研究中,当通过AUC、FDR和MDR在一系列模拟参数范围内进行评估时,贝叶斯模型通常比其他方法更好,并且在个别研究中,这种差异随着样本量的增加而减小。分裂研究验证表明,在没有平台差异、样本差异和注释差异的情况下,贝叶斯模型的适当收缩,否则会使实验数据分析复杂化。最后,我们将我们的模型适用于四项乳腺癌研究,这些研究使用不同的技术(cDNA和Affymetrix)来估计雌激素受体阳性肿瘤与雌激素受体阴性肿瘤的差异表达。用于复制我们的分析的软件和数据是公开可用的。
In this paper we define a hierarchical Bayesian model for microarray expression data collected from several studies and use it to identify genes that show differential expression between two conditions. Key features include shrinkage across both genes and studies, and flexible modeling that allows for interactions between platforms and the estimated effect, as well as concordant and discordant differential expression across studies. We evaluated the performance of our model in a comprehensive fashion, using both artificial data, and a “split-study” validation approach that provides an agnostic assessment of the model's behavior not only under the null hypothesis, but also under a realistic alternative. The simulation results from the artificial data demonstrate the advantages of the Bayesian model. The 1 – AUC values for the Bayesian model are roughly half of the corresponding values for a direct combination of t- and SAM-statistics. Furthermore, the simulations provide guidelines for when the Bayesian model is most likely to be useful. Most noticeably, in small studies the Bayesian model generally outperforms other methods when evaluated by AUC, FDR, and MDR across a range of simulation parameters, and this difference diminishes for larger sample sizes in the individual studies. The split-study validation illustrates appropriate shrinkage of the Bayesian model in the absence of platform-, sample-, and annotation-differences that otherwise complicate experimental data analyses. Finally, we fit our model to four breast cancer studies employing different technologies (cDNA and Affymetrix) to estimate differential expression in estrogen receptor positive tumors versus negative ones. Software and data for reproducing our analysis are publicly available.
DOI: 10.1186/gb-2004-5-10-r80
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Gentleman RC;Carey VJ;Bates DM;Bolstad B;Dettling M;Dudoit S;Ellis B;Gautier L;Ge Y;Gentry J;Hornik K;Hothorn T;Huber W;Iacus S;Irizarry R;Leisch F;Li C;Maechler M;Rossini AJ;Sawitzki G;Smith C;Smyth G;Tierney L;Yang JY;Zhang J
通讯作者: Zhang J
DOI: 10.1186/1471-2105-7-247
发表时间: 2006-05-05
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Conlon, Erin M.;Song, Joon J.;Liu, Jun S.
通讯作者: Liu, Jun S.
DOI: 10.1093/biostatistics/kxm033
发表时间: 2008-04-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Garrett-Mayer, Elizabeth;Parmigiani, Giovanni;Gabrielson, Edward
通讯作者: Gabrielson, Edward
DOI: 10.1038/sj.onc.1208561
发表时间: 2005-07-01
期刊: ONCOGENE
影响因子: 8
作者:
Farmer, P;Bonnefoi, H;Iggo, R
通讯作者: Iggo, R
来自多个微阵列实验的基因表达数据的荟萃分析的潜在变量方法。
DOI: 10.1186/1471-2105-8-364
发表时间: 2007-09-27
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Choi, Hyungwon;Shen, Ronglai;Chinnaiyan, Arul M;Ghosh, Debashis
通讯作者: Ghosh, Debashis