High resolution genome wide binding event finding and motif discovery reveals transcription factor spatial binding constraints.

High resolution genome wide binding event finding and motif discovery reveals transcription factor spatial binding constraints.
复制标题

DOI:
10.1371/journal.pcbi.1002638
复制
发表时间:
2012
影响因子:
4.3
通讯作者:
Gifford DK
Gifford DK
中科院分区:
生物学2区
文献类型:
--
作者:
Guo Y;Mahony S;Gifford DK

文献摘要

参考文献

被引文献

相似文献

基因组功能的一个重要组成部分是基因组调控元件的语法,它决定了不同的转录因子如何相互作用以编排一个调控控制程序。对关键转录因子之间体内间距限制的精确描述将揭示这种基因组调控语言的关键方面。为了在体内发现新的转录因子空间结合限制,我们开发了一种新的综合计算方法,即全基因组事件发现和基序发现(GEM)。GEM通过在染色质免疫沉淀(ChIP)数据和基因组序列的生成概率模型背景下,将结合事件发现和基序发现与位置先验联系起来,以高空间分辨率将ChIP数据解析为解释性基序和结合事件。对214个ENCODE人类ChIP - Seq实验中的63个转录因子进行GEM分析,比其他当代方法恢复了更多已知的因子基序,并为结合特异性未知的因子发现了6个新基序。GEM对结合事件读取分布的自适应学习使其能够进一步改进先前处理ChIP - Seq和ChIP - exo数据的方法,以产生无与伦比的空间分辨率,并发现同一因子紧密间隔的结合事件。在使用GEM对体内序列特异性转录因子结合进行的系统分析中,我们发现了因子之间数百种空间结合限制。GEM在小鼠胚胎干细胞中发现了37个因子结合限制的例子,包括Klf4与其他关键调控因子之间强烈的距离特异性限制。在人类ENCODE数据中,GEM发现了390个空间受限的成对结合的例子,包括诸如c - Fos:c - Jun/USF1、CTCF/Egr1和HNF4A/FOXA1等新的成对组合。在ChIP数据中发现新的因子 - 因子空间限制是重要的,因为它为调控因子相互作用提出了可测试的模型,这将有助于阐明基因组功能以及组合控制的实现。 我们基因组中的字母组成单词和短语,它们控制着每个基因何时被激活。为了理解这些单词和短语在健康和疾病中是如何起作用的,我们开发了一种新的计算方法,以确定每个基因组调控蛋白在我们的基因组文本中使用哪些单词位置,以及这些活性单词彼此之间是如何间隔的。我们的方法通过将实验数据与我们的基因组文本相结合,以找到受每个蛋白因子调控的精确单词,从而实现了卓越的空间准确性。通过这种分析,我们在实验数据中发现了新的单词间距,这暗示了新的基因组语法控制结构。
An essential component of genome function is the syntax of genomic regulatory elements that determine how diverse transcription factors interact to orchestrate a program of regulatory control. A precise characterization of in vivo spacing constraints between key transcription factors would reveal key aspects of this genomic regulatory language. To discover novel transcription factor spatial binding constraints in vivo, we developed a new integrative computational method, genome wide event finding and motif discovery (GEM). GEM resolves ChIP data into explanatory motifs and binding events at high spatial resolution by linking binding event discovery and motif discovery with positional priors in the context of a generative probabilistic model of ChIP data and genome sequence. GEM analysis of 63 transcription factors in 214 ENCODE human ChIP-Seq experiments recovers more known factor motifs than other contemporary methods, and discovers six new motifs for factors with unknown binding specificity. GEM's adaptive learning of binding-event read distributions allows it to further improve upon previous methods for processing ChIP-Seq and ChIP-exo data to yield unsurpassed spatial resolution and discovery of closely spaced binding events of the same factor. In a systematic analysis of in vivo sequence-specific transcription factor binding using GEM, we have found hundreds of spatial binding constraints between factors. GEM found 37 examples of factor binding constraints in mouse ES cells, including strong distance-specific constraints between Klf4 and other key regulatory factors. In human ENCODE data, GEM found 390 examples of spatially constrained pair-wise binding, including such novel pairs as c-Fos:c-Jun/USF1, CTCF/Egr1, and HNF4A/FOXA1. The discovery of new factor-factor spatial constraints in ChIP data is significant because it proposes testable models for regulatory factor interactions that will help elucidate genome function and the implementation of combinatorial control. The letters in our genome spell words and phrases that control when each gene is activated. To understand how these words and phrases function in health and disease, we have developed a new computational method to determine what word positions in our genomic text are used by each genome regulatory protein, and how these active words are spaced relative to one another. Our method achieves exceptional spatial accuracy by integrating experimental data with the text of our genome to find the precise words that are regulated by each protein factor. Using this analysis we have discovered novel word spacings in the experimental data that suggest novel genome grammatical control constructs.
DOI: 10.1038/nbt.1505
发表时间: 2008-11
影响因子: 46.9
作者:
Ji, Hongkai;Jiang, Hui;Ma, Wenxiu;Johnson, David S.;Myers, Richard M.;Wong, Wing H.
通讯作者: Wong, Wing H.
DOI: 10.1186/1471-2105-12-139
发表时间: 2011-05-09
期刊: BMC bioinformatics
影响因子: 3
作者:
Feng X;Grossman R;Stein L
通讯作者: Stein L
DOI: 10.1016/j.stem.2009.12.009
发表时间: 2010-02-05
期刊: CELL STEM CELL
影响因子: 23.9
作者:
Heng, Jian-Chien Dominic;Feng, Bo;Ng, Huck-Hui
通讯作者: Ng, Huck-Hui
DOI: 10.1038/sj.onc.1205400
发表时间: 2002-05-13
期刊: ONCOGENE
影响因子: 8
作者:
Hoffman, B;Amanullah, A;Liebermann, DA
通讯作者: Liebermann, DA
DOI: 10.1038/nbt1246
发表时间: 2006-11-01
影响因子: 46.9
作者:
Berger, Michael F.;Philippakis, Anthony A.;Bulyk, Martha L.
通讯作者: Bulyk, Martha L.