A Bayesian model selection approach for identifying differentially expressed transcripts from RNA sequencing data.

A Bayesian model selection approach for identifying differentially expressed transcripts from RNA sequencing data.
复制标题

DOI:
10.1111/rssc.12213
复制
发表时间:
2018-01
期刊:
Journal of the Royal Statistical Society. Series C, Applied statistics
影响因子:
--
通讯作者:
Rattray M
Rattray M
中科院分区:
其他
文献类型:
--
作者:
Papastamoulis P;Rattray M

文献摘要

相似文献

分子生物学的最新进展允许量化转录组并将转录物评分为在两种生物条件之间差异表达或相等表达。虽然这两个任务是密切相关的,但可用的推理方法将它们分开处理:使用主模型来估计表达,并使用差分表达模型对其输出进行后处理。在本文中,这两个问题同时解决,提出联合估计的表达水平和差异表达:未知的相对丰度的每个转录本可以是相等的或不之间的两个条件。在BitSeq框架上建立了一个分层贝叶斯模型,并通过使用马尔可夫链蒙特卡罗抽样来推断转录表达和差异表达的后验分布。结果表明,所提出的模型具有共轭的固定维变量,从而完整的条件分布解析推导。构造了两个采样器,一个可逆跳马尔可夫链蒙特卡罗采样器和一个崩溃的吉布斯采样器,后者被发现表现得更好。一个集群表示的对齐读取的转录组的介绍,允许在合理的计算时间下的转录本的子集的边缘后验分布的并行估计。在微分表达式的先验概率固定的情况下,聚类采样器具有与原始采样器相同的边缘后验分布,但也采用了更一般的先验结构。提出的算法是基准对替代方法,通过使用合成的数据集和应用到真实的RNA测序数据。源代码可从https://github.com/mqbssppe/cjBitSeq在线获得。
Recent advances in molecular biology allow the quantification of the transcriptome and scoring transcripts as differentially or equally expressed between two biological conditions. Although these two tasks are closely linked, the available inference methods treat them separately: a primary model is used to estimate expression and its output is post processed by using a differential expression model. In the paper, both issues are simultaneously addressed by proposing the joint estimation of expression levels and differential expression: the unknown relative abundance of each transcript can either be equal or not between two conditions. A hierarchical Bayesian model builds on the BitSeq framework and the posterior distribution of transcript expression and differential expression is inferred by using Markov chain Monte Carlo sampling. It is shown that the model proposed enjoys conjugacy for fixed dimension variables; thus the full conditional distributions are analytically derived. Two samplers are constructed, a reversible jump Markov chain Monte Carlo sampler and a collapsed Gibbs sampler, and the latter is found to perform better. A cluster representation of the aligned reads to the transcriptome is introduced, allowing parallel estimation of the marginal posterior distribution of subsets of transcripts under reasonable computing time. Under a fixed prior probability of differential expression the clusterwise sampler has the same marginal posterior distributions as the raw sampler, but a more general prior structure is also employed. The algorithm proposed is benchmarked against alternative methods by using synthetic data sets and applied to real RNA sequencing data. Source code is available on line from https://github.com/mqbssppe/cjBitSeq.