Efficient inference for sparse latent variable models of transcriptional regulation.

Efficient inference for sparse latent variable models of transcriptional regulation.
复制标题

DOI:
10.1093/bioinformatics/btx508
复制
发表时间:
2017-12-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Rattray M
Rattray M
中科院分区:
其他
文献类型:
--
作者:
Dai Z;Iqbal M;Lawrence ND;Rattray M

文献摘要

参考文献

相似文献

原核生物基因表达的调控涉及复杂的共调控机制,涉及大量的转录调控蛋白及其靶基因。揭示这些基因组规模的相互作用构成了系统生物学的主要瓶颈。稀疏潜在因子模型,假设转录因子(TF)的活性为未观察到的,提供了一个生物学上可解释的建模框架,整合基因表达和全基因组结合数据,但在同一时间提出了一个困难的计算推理问题。现有的这种模型的概率推理方法依赖于主观过滤,并遭受可扩展性问题,因此不适合现实的基因组规模的应用。我们提出了一个快速贝叶斯稀疏因子模型,该模型采用来自ChIP-seq实验或基序预测的输入基因表达和结合位点数据,并输出活性TF-基因链接以及潜在TF活性。我们的方法采用了一种有效的变分贝叶斯模型推理方案,使其应用于大型数据集,这是不可行的,现有的MCMC为基础的推理方法,这样的模型。我们验证了我们的方法对类似的模型在文献中的合成数据,采用MCMC推理,并获得可比的结果与一小部分的计算时间。我们还将我们的方法应用于结核分枝杆菌的大规模数据,涉及113个TF的ChIP-seq数据和3863个推定靶基因的匹配基因表达数据。我们评估我们的预测使用一个独立的转录组学实验涉及过度表达的TF。我们的方法与数据的一个易于使用的笔记本电脑演示可在https://github.com/zhenwendai/SITAR。 补充数据可在Bioinformatics在线获得。
Regulation of gene expression in prokaryotes involves complex co-regulatory mechanisms involving large numbers of transcriptional regulatory proteins and their target genes. Uncovering these genome-scale interactions constitutes a major bottleneck in systems biology. Sparse latent factor models, assuming activity of transcription factors (TFs) as unobserved, provide a biologically interpretable modelling framework, integrating gene expression and genome-wide binding data, but at the same time pose a hard computational inference problem. Existing probabilistic inference methods for such models rely on subjective filtering and suffer from scalability issues, thus are not well-suited for realistic genome-scale applications. We present a fast Bayesian sparse factor model, which takes input gene expression and binding sites data, either from ChIP-seq experiments or motif predictions, and outputs active TF-gene links as well as latent TF activities. Our method employs an efficient variational Bayes scheme for model inference enabling its application to large datasets which was not feasible with existing MCMC-based inference methods for such models. We validate our method on synthetic data against a similar model in the literature, employing MCMC for inference, and obtain comparable results with a small fraction of the computational time. We also apply our method to large-scale data from Mycobacterium tuberculosis involving ChIP-seq data on 113 TFs and matched gene expression data for 3863 putative target genes. We evaluate our predictions using an independent transcriptomics experiment involving over-expression of TFs. An easy-to-use Jupyter notebook demo of our method with data is available at https://github.com/zhenwendai/SITAR. Supplementary data are available at Bioinformatics online.
DOI: 10.15252/msb.20156236
发表时间: 2015-11-17
影响因子: 9.9
作者:
Arrieta-Ortiz ML;Hafemeister C;Bate AR;Chu T;Greenfield A;Shuster B;Barry SN;Gallitto M;Liu B;Kacmarczyk T;Santoriello F;Chen J;Rodrigues CD;Sato T;Rudner DZ;Driks A;Bonneau R;Eichenberger P
通讯作者: Eichenberger P
DOI: 10.1073/pnas.0913357107
发表时间: 2010-04-06
影响因子: 11.1
作者:
Marbach, Daniel;Prill, Robert J.;Stolovitzky, Gustavo
通讯作者: Stolovitzky, Gustavo
DOI: 10.1073/pnas.112341999
发表时间: 2002-09-03
影响因子: 11.1
作者:
Li, H;Rhodius, V;Siggia, ED
通讯作者: Siggia, ED
DOI: 10.1073/pnas.2136632100
发表时间: 2003-12-23
影响因子: 11.1
作者:
Liao, JC;Boscolo, R;Roychowdhury, VP
通讯作者: Roychowdhury, VP
DOI: 10.1038/nmeth.2016
发表时间: 2012-07-15
期刊: NATURE METHODS
影响因子: 48
作者:
Marbach, Daniel;Costello, James C.;Kueffner, Robert;Vega, Nicole M.;Prill, Robert J.;Camacho, Diogo M.;Allison, Kyle R.;Kellis, Manolis;Collins, James J.;Stolovitzky, Gustavo
通讯作者: Stolovitzky, Gustavo