On the value of intra-motif dependencies of human insulator protein CTCF.

On the value of intra-motif dependencies of human insulator protein CTCF.
复制标题

DOI:
10.1371/journal.pone.0085629
复制
发表时间:
2014
期刊:
影响因子:
3.7
通讯作者:
Grosse I
Grosse I
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Eggeling R;Gohr A;Keilwagen J;Mohr M;Posch S;Smith AD;Grosse I

文献摘要

参考文献

被引文献

相似文献

DNA结合蛋白如转录因子的结合亲和力主要由DNA链上相应结合位点的碱基组成决定。大多数蛋白质不仅结合单个序列,而是结合一组序列,这些序列可以通过序列基序建模。从头发现基序的算法在启动子模型、学习方法和其他方面有所不同,但通常使用统计上简单的基序位置权重矩阵模型,该模型假设所有核苷酸之间的统计独立性。然而,这种假设没有明确的理由,导致关于对结合位点内核苷酸之间的依赖性进行建模的重要性的持续争论。过去,对结合位点内的统计依赖性进行建模一直受到数据有限问题的阻碍。随着 ChIP-seq 等高通量技术的兴起,这种情况现在已经改变,使得有效利用统计依赖性成为可能。在这项工作中,我们通过使用最近开发的非均匀简约马尔可夫模型的模型类来研究人类增强子阻断绝缘体蛋白 CTCF 结合位点中统计依赖性的存在,该模型能够对复杂的依赖性进行建模,同时避免过度拟合。这些发现导致了 CTCF 结合基序的更详细表征,该基序仅通过几个位置(主要是 3' 端)的独立核苷酸频率来表示。
The binding affinity of DNA-binding proteins such as transcription factors is mainly determined by the base composition of the corresponding binding site on the DNA strand. Most proteins do not bind only a single sequence, but rather a set of sequences, which may be modeled by a sequence motif. Algorithms for de novo motif discovery differ in their promoter models, learning approaches, and other aspects, but typically use the statistically simple position weight matrix model for the motif, which assumes statistical independence among all nucleotides. However, there is no clear justification for that assumption, leading to an ongoing debate about the importance of modeling dependencies between nucleotides within binding sites. In the past, modeling statistical dependencies within binding sites has been hampered by the problem of limited data. With the rise of high-throughput technologies such as ChIP-seq, this situation has now changed, making it possible to make use of statistical dependencies effectively. In this work, we investigate the presence of statistical dependencies in binding sites of the human enhancer-blocking insulator protein CTCF by using the recently developed model class of inhomogeneous parsimonious Markov models, which is capable of modeling complex dependencies while avoiding overfitting. These findings lead to a more detailed characterization of the CTCF binding motif, which is only poorly represented by independent nucleotide frequencies at several positions, predominantly at the 3′ end.
DOI: 10.1371/journal.pcbi.1001070
发表时间: 2011-02-10
影响因子: 4.3
作者:
Keilwagen J;Grau J;Paponov IA;Posch S;Strickert M;Grosse I
通讯作者: Grosse I
DOI: 10.1126/science.1162327
发表时间: 2009-06-26
期刊: Science (New York, N.Y.)
影响因子: --
作者:
Badis G;Berger MF;Philippakis AA;Talukder S;Gehrke AR;Jaeger SA;Chan ET;Metzler G;Vedenko A;Chen X;Kuznetsov H;Wang CF;Coburn D;Newburger DE;Morris Q;Hughes TR;Bulyk ML
通讯作者: Bulyk ML
DOI: 10.1038/nbt1246
发表时间: 2006-11-01
影响因子: 46.9
作者:
Berger, Michael F.;Philippakis, Anthony A.;Bulyk, Martha L.
通讯作者: Bulyk, Martha L.
DOI: 10.1093/nar/30.5.1255
发表时间: 2002-03-01
影响因子: 14.9
作者:
Bulyk, ML;Johnson, PLF;Church, GM
通讯作者: Church, GM
DOI: 10.1016/j.cell.2006.12.048
发表时间: 2007-03-23
期刊: CELL
影响因子: 64.5
作者:
Kim, Tae Hoon;Abdullaev, Ziedulla K.;Ren, Bing
通讯作者: Ren, Bing