Modeling allele-specific expression at the gene and SNP levels simultaneously by a Bayesian logistic mixed regression model

Modeling allele-specific expression at the gene and SNP levels simultaneously by a Bayesian logistic mixed regression model
复制标题

DOI:
10.1186/s12859-019-3141-6
复制
发表时间:
2019-10-28
期刊:
影响因子:
3
通讯作者:
Rivera, Rocio M.
Rivera, Rocio M.
中科院分区:
生物学4区
文献类型:
--
作者:
Xie, Jing;Ji, Tieming;Rivera, Rocio M.

文献摘要

被引文献

相似文献

背景:可以确定等位基因来源的高通量测序实验已被用于评估全基因组的等位基因特异性表达。尽管高通量实验产生了大量数据,但统计方法往往过于简单化,无法理解基因表达的复杂性。具体地说,现有的方法没有单独和同时测试一个基因作为一个整体的等位基因特异性表达(ASE)和基因内ASE在外显子之间的变异。结果:我们提出了一个广义线性混合模型来弥合这些差距,将基因、单核苷酸多态(SNPs)和生物复制的变异纳入其中。为了提高统计推断的可靠性,我们为模型中的每一种效应分配先验,以便信息在整个基因组中的基因之间共享。我们利用贝叶斯模型选择来检验每个基因的ASE假设以及基因内SNPs之间的变异。我们将我们的方法应用于牛研究中的四种组织类型,以从头检测牛基因组中的ASE基因,并发现跨基因外显子和跨组织类型的调控ASE的有趣预测。我们通过模仿真实数据集的模拟研究,将我们的方法与竞争对手的方法进行了比较。实现我们提出的算法的R包BLMRM可以在https://github.com/JingXieMIZZOU/BLMRM.Conclusions:上下载。我们将证明,当存在SNP变异和生物变异时,所提出的方法具有更好的误发现率控制和比现有方法更好的能力。此外,我们的方法还保持了较低的计算要求,允许进行全基因组分析。
Background: High-throughput sequencing experiments, which can determine allele origins, have been used to assess genome-wide allele-specific expression. Despite the amount of data generated from high-throughput experiments, statistical methods are often too simplistic to understand the complexity of gene expression. Specifically, existing methods do not test allele-specific expression (ASE) of a gene as a whole and variation in ASE within a gene across exons separately and simultaneously.Results: We propose a generalized linear mixed model to close these gaps, incorporating variations due to genes, single nucleotide polymorphisms (SNPs), and biological replicates. To improve reliability of statistical inferences, we assign priors on each effect in the model so that information is shared across genes in the entire genome. We utilize Bayesian model selection to test the hypothesis of ASE for each gene and variations across SNPs within a gene. We apply our method to four tissue types in a bovine study to de novo detect ASE genes in the bovine genome, and uncover intriguing predictions of regulatory ASEs across gene exons and across tissue types. We compared our method to competing approaches through simulation studies that mimicked the real datasets. The R package, BLMRM, that implements our proposed algorithm, is publicly available for download at https://github.com/JingXieMIZZOU/BLMRM.Conclusions: We will show that the proposed method exhibits improved control of the false discovery rate and improved power over existing methods when SNP variation and biological variation are present. Besides, our method also maintains low computational requirements that allows for whole genome analysis.