Empirical Bayesian analysis of paired high-throughput sequencing data with a beta-binomial distribution.

Empirical Bayesian analysis of paired high-throughput sequencing data with a beta-binomial distribution.
复制标题

DOI:
10.1186/1471-2105-14-135
复制
发表时间:
2013-04-23
期刊:
影响因子:
3
通讯作者:
Kelly KA
Kelly KA
中科院分区:
生物学4区
文献类型:
--
作者:
Hardcastle TJ;Kelly KA

文献摘要

参考文献

被引文献

相似文献

在许多基因组实验中,样本配对是自然发生的。例如,同一患者的肿瘤和正常组织中的基因表达。需要分析此类实验的高通量测序数据的方法来识别不同实验条件下配对样品内和配对之间的差异表达。我们开发了一种基于 β 二项式分布的经验贝叶斯方法,用于对高通量测序实验中的配对数据进行建模。我们在各种场景中检查该方法在模拟和真实数据上的性能。我们的方法是作为 Bioconductor (http://www.bioconductor.org) 提供的 RbaySeq 包(版本 1.11.6 及更高版本)的一部分实现的。我们将我们的方法与基于广义线性建模方法的替代方法进行比较,并表明我们的方法在模拟数据的性能方面提供了显着的提升。在对口腔鳞状细胞癌患者的真实数据进行测试时,我们发现先前确定的头颈鳞状细胞癌相关基因集比之前通过广义线性建模方法实现的基因集更加丰富,这表明在真实数据中可能会发现类似的性能增益。因此,我们的方法显示了对配对样本的高通量测序数据分析的真正和实质性的改进。
Pairing of samples arises naturally in many genomic experiments; for example, gene expression in tumour and normal tissue from the same patients. Methods for analysing high-throughput sequencing data from such experiments are required to identify differential expression, both within paired samples and between pairs under different experimental conditions. We develop an empirical Bayesian method based on the beta-binomial distribution to model paired data from high-throughput sequencing experiments. We examine the performance of this method on simulated and real data in a variety of scenarios. Our methods are implemented as part of the RbaySeq package (versions 1.11.6 and greater) available from Bioconductor (http://www.bioconductor.org). We compare our approach to alternatives based on generalised linear modelling approaches and show that our method offers significant gains in performance on simulated data. In testing on real data from oral squamous cell carcinoma patients, we discover greater enrichment of previously identified head and neck squamous cell carcinoma associated gene sets than has previously been achieved through a generalised linear modelling approach, suggesting that similar gains in performance may be found in real data. Our methods thus show real and substantial improvements in analyses of high-throughput sequencing data from paired samples.
DOI: 10.1371/journal.pone.0003215
发表时间: 2008-09-15
期刊: PLOS ONE
影响因子: 3.7
作者:
Yu, Yau-Hua;Kuo, Hsu-Ko;Chang, Kuo-Wei
通讯作者: Chang, Kuo-Wei
DOI: 10.1186/gb-2010-11-3-r25
发表时间: 2010
期刊: Genome biology
影响因子: 12.3
作者:
Robinson MD;Oshlack A
通讯作者: Oshlack A
DOI: 10.1158/0008-5472.can-06-3322
发表时间: 2007-04-01
期刊: CANCER RESEARCH
影响因子: 11.2
作者:
Winter, Stuart C.;Buffa, Francesca M.;Harris, Adrian L.
通讯作者: Harris, Adrian L.
DOI: 10.1093/nar/gks042
发表时间: 2012-05
影响因子: 14.9
作者:
McCarthy DJ;Chen Y;Smyth GK
通讯作者: Smyth GK
DOI: 10.1371/journal.pone.0031630
发表时间: 2012
期刊: PloS one
影响因子: 3.7
作者:
Cordero F;Beccuti M;Arigoni M;Donatelli S;Calogero RA
通讯作者: Calogero RA