Bayesian analysis of RNA sequencing data by estimating multiple shrinkage priors

Bayesian analysis of RNA sequencing data by estimating multiple shrinkage priors
复制标题

DOI:
10.1093/biostatistics/kxs031
复制
发表时间:
2013-01-01
期刊:
影响因子:
2.1
通讯作者:
Van Wieringen, Wessel N.
Van Wieringen, Wessel N.
中科院分区:
数学2区
文献类型:
--
作者:
Van De Wiel, Mark A.;Leday, Gwenael G. R.;Van Wieringen, Wessel N.

文献摘要

被引文献

相似文献

下一代测序正在迅速取代微阵列作为探测细胞不同分子水平(例如 DNA 或 RNA)的技术。该技术提供了更高的分辨率,同时减少了偏差。 RNA 测序结果是 RNA 链的计数。此类数据带来了新的统计挑战。我们提出了一种新颖的通用方法来建模和分析此类数据。我们的方法旨在使似然(计数)模型和回归模型具有更大的灵活性。因此,支持多种计数模型,例如流行的 NB 模型,它解释了过度分散的问题。此外,还适应复杂的、非平衡的设计和随机效应。与其他一些方法一样,我们的方法提供了色散相关参数的收缩。然而,我们通过启用参数的联合收缩来扩展它,包括那些需要推理的参数。我们认为这对于贝叶斯多重性校正至关重要。收缩是通过经验估计先验来实现的。我们讨论了几个参数(混合)和非参数先验,并开发了估计这些先验(参数)的程序。通过本地和贝叶斯错误发现率提供推断。我们在几次模拟和两个数据集上说明了我们的方法,并将其与其他方法进行了比较。基于模型和数据的模拟显示在给定特异性下灵敏度有了显着提高。这些数据促使人们使用 ZI-NB 作为 NB 的强大替代品,从而提高低计数数据的检测率。最后,与其他方法相比,在大样本补充上进行验证时,小样本子集的结果更具可重复性,这说明了收缩类型的重要性。
Next generation sequencing is quickly replacing microarrays as a technique to probe different molecular levels of the cell, such as DNA or RNA. The technology provides higher resolution, while reducing bias. RNA sequencing results in counts of RNA strands. This type of data imposes new statistical challenges. We present a novel, generic approach to model and analyze such data. Our approach aims at large flexibility of the likelihood (count) model and the regression model alike. Hence, a variety of count models is supported, such as the popular NB model, which accounts for overdispersion. In addition, complex, non-balanced designs and random effects are accommodated. Like some other methods, our method provides shrinkage of dispersion-related parameters. However, we extend it by enabling joint shrinkage of parameters, including those for which inference is desired. We argue that this is essential for Bayesian multiplicity correction. Shrinkage is effectuated by empirically estimating priors. We discuss several parametric (mixture) and non-parametric priors and develop procedures to estimate (parameters of) those. Inference is provided by means of local and Bayesian false discovery rates. We illustrate our method on several simulations and two data sets, also to compare it with other methods. Model- and data-based simulations show substantial improvements in the sensitivity at the given specificity. The data motivate the use of the ZI-NB as a powerful alternative to the NB, which results in higher detection rates for low-count data. Finally, compared with other methods, the results on small sample subsets are more reproducible when validated on their large sample complements, illustrating the importance of the type of shrinkage.