Probabilistic inference of transcription factor binding from multiple data sources.

Probabilistic inference of transcription factor binding from multiple data sources.
复制标题

DOI:
10.1371/journal.pone.0001820
复制
发表时间:
2008-03-26
期刊:
影响因子:
3.7
通讯作者:
Shmulevich I
Shmulevich I
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Lähdesmäki H;Rust AG;Shmulevich I

文献摘要

参考文献

被引文献

相似文献

分子生物学中的一个重要问题是建立对细胞中转录调控过程的完整理解。我们已经开发了一个灵活的概率框架,从多个数据源预测TF结合,在几个方面不同于标准的假设检验(扫描)方法。我们的概率建模框架估计绑定的概率,因此,自然反映了我们对绑定的信念程度。概率建模还允许将我们的结合预测简单而系统地整合到其他概率建模方法中,例如基于表达的基因网络推断。该方法回答了整个分析的启动子是否具有结合位点的问题,但也可以扩展到估计每个核苷酸位置处的结合概率。此外,我们引入了一个扩展模型组合监管的几个TF。最重要的是,所提出的方法可以从多个证据来源,如多个统计模型(基序)的TF,进化保守性,调节潜力,CpG岛,核小体定位,DNA酶超敏位点,ChIP芯片结合片段和其他(先验)基于序列的生物学知识进行原则性概率推断。我们开发了一种可能性和贝叶斯方法,后者是用马尔可夫链蒙特卡罗算法实现的。从小鼠基因组中精心构建的测试集的结果表明,原则性数据融合可以显着提高TF结合预测方法的性能。我们还将概率建模框架应用于小鼠基因组中的所有启动子,结果表明转录调控因子与其靶启动子之间存在稀疏连接。为了便于分析其他序列和额外的数据,我们已经开发了一个在线的网络工具,ProbTF,它实现了我们的概率TF结合预测方法,使用多个数据源。测试数据集、网络工具、源代码和补充数据可在http://www.probtf.org上获得。
An important problem in molecular biology is to build a complete understanding of transcriptional regulatory processes in the cell. We have developed a flexible, probabilistic framework to predict TF binding from multiple data sources that differs from the standard hypothesis testing (scanning) methods in several ways. Our probabilistic modeling framework estimates the probability of binding and, thus, naturally reflects our degree of belief in binding. Probabilistic modeling also allows for easy and systematic integration of our binding predictions into other probabilistic modeling methods, such as expression-based gene network inference. The method answers the question of whether the whole analyzed promoter has a binding site, but can also be extended to estimate the binding probability at each nucleotide position. Further, we introduce an extension to model combinatorial regulation by several TFs. Most importantly, the proposed methods can make principled probabilistic inference from multiple evidence sources, such as, multiple statistical models (motifs) of the TFs, evolutionary conservation, regulatory potential, CpG islands, nucleosome positioning, DNase hypersensitive sites, ChIP-chip binding segments and other (prior) sequence-based biological knowledge. We developed both a likelihood and a Bayesian method, where the latter is implemented with a Markov chain Monte Carlo algorithm. Results on a carefully constructed test set from the mouse genome demonstrate that principled data fusion can significantly improve the performance of TF binding prediction methods. We also applied the probabilistic modeling framework to all promoters in the mouse genome and the results indicate a sparse connectivity between transcriptional regulators and their target promoters. To facilitate analysis of other sequences and additional data, we have developed an on-line web tool, ProbTF, which implements our probabilistic TF binding prediction method using multiple data sources. Test data set, a web tool, source codes and supplementary data are available at: http://www.probtf.org.
DOI: 10.1016/s0092-8674(04)00127-8
发表时间: 2004-02-20
期刊: CELL
影响因子: 64.5
作者:
Cawley, S;Bekiranov, S;Gingeras, TR
通讯作者: Gingeras, TR
DOI: 10.1371/journal.pcbi.0020070
发表时间: 2006-06-16
影响因子: 4.3
作者:
Beyer A;Workman C;Hollunder J;Radke D;Möller U;Wilhelm T;Ideker T
通讯作者: Ideker T
DOI: 10.1089/cmb.2005.12.314
发表时间: 2005-04-01
影响因子: 1.7
作者:
Hertzberg, L;Zuk, O;Domany, E
通讯作者: Domany, E
DOI: 10.1093/nar/gkj116
发表时间: 2006-01-01
影响因子: 14.9
作者:
Blanco E;Farré D;Albà MM;Messeguer X;Guigó R
通讯作者: Guigó R
DOI: 10.1016/j.cell.2005.10.042
发表时间: 2006-01-13
期刊: CELL
影响因子: 64.5
作者:
Hallikas, O;Palin, K;Taipale, J
通讯作者: Taipale, J