Statistical modeling of transcription factor binding affinities predicts regulatory interactions.

Statistical modeling of transcription factor binding affinities predicts regulatory interactions.
复制标题

DOI:
10.1371/journal.pcbi.1000039
复制
发表时间:
2008-03-21
影响因子:
4.3
通讯作者:
Vingron M
Vingron M
中科院分区:
生物学2区
文献类型:
--
作者:
Manke T;Roider HG;Vingron M

文献摘要

参考文献

被引文献

相似文献

最近的实验和理论工作已经强调了这样一个事实,即转录因子与DNA的结合可以通过连续测量其结合亲和力而不是结合位点的离散描述来更准确地描述。虽然结合亲和力可以从物理模型预测,但通常希望知道特定序列背景的结合亲和力的分布。在本文中,我们提出了一种统计方法来获得精确的分布与固定GC含量的序列模型。我们证明,几乎所有已知的转录因子的亲和力分布可以有效地参数化一类广义极值分布。此外,该参数化还描述了具有可变GC含量的序列背景(例如人启动子序列)的亲和力分布。我们的方法适用于任意序列和所有已知的结合偏好,可以描述的基序矩阵的转录因子。统计处理还提供了一个适当的框架,直接比较具有非常不同的亲和力分布的转录因子。我们对具有已知结合位点的人类启动子的分析说明了这一点,对于其中的许多,我们可以将已知的调节剂鉴定为具有最高亲和力的那些。物理模型和统计标准化的组合提供了一种定量测量,其对给定序列的转录因子进行排名,并且可以直接与大规模结合数据进行比较。它的成功应用于人类启动子序列作为一个令人鼓舞的例子,该方法可以应用于其他序列。蛋白质与DNA的结合是一种关键的分子机制,它可以调节基因的表达以响应不同的细胞和环境条件。对基因调控的广泛研究已经产生了许多转录因子的结合模型,但是新结合位点的预测仍然具有挑战性,并且难以以任何系统的方式改进。最近的实验进展,特别是高通量结合测定,已经将理论重点从预测新的结合位点转移到转录因子的结合亲和力的更定量模型,其现在可以在整个基因组中测量。因此,我们开发了一个生物物理模型,该模型解释了所观察到的结合强度的变化。在这里,我们扩展这个框架模型不仅结合亲和力,而且其分布在不同的序列背景。这使我们能够比较来自不同转录因子的预测亲和力,并根据其归一化亲和力对其进行排名。这种排名的生物学意义是什么?我们已经证明,许多已知的转录因子和它们各自的目标之间的关联表现为强相互作用。这提供了一个基本原理来预测,对于任何给定的启动子区域,那些转录因子是最有可能参与其调节。
Recent experimental and theoretical efforts have highlighted the fact that binding of transcription factors to DNA can be more accurately described by continuous measures of their binding affinities, rather than a discrete description in terms of binding sites. While the binding affinities can be predicted from a physical model, it is often desirable to know the distribution of binding affinities for specific sequence backgrounds. In this paper, we present a statistical approach to derive the exact distribution for sequence models with fixed GC content. We demonstrate that the affinity distribution of almost all known transcription factors can be effectively parametrized by a class of generalized extreme value distributions. Moreover, this parameterization also describes the affinity distribution for sequence backgrounds with variable GC content, such as human promoter sequences. Our approach is applicable to arbitrary sequences and all transcription factors with known binding preferences that can be described in terms of a motif matrix. The statistical treatment also provides a proper framework to directly compare transcription factors with very different affinity distributions. This is illustrated by our analysis of human promoters with known binding sites, for many of which we could identify the known regulators as those with the highest affinity. The combination of physical model and statistical normalization provides a quantitative measure which ranks transcription factors for a given sequence, and which can be compared directly with large-scale binding data. Its successful application to human promoter sequences serves as an encouraging example of how the method can be applied to other sequences. The binding of proteins to DNA is a key molecular mechanism, which can regulate the expression of genes in response to different cellular and environmental conditions. The extensive research on gene regulation has generated binding models for many transcription factors, but the prediction of new binding sites is still challenging and difficult to improve in any systematic way. Recent experimental advances, notably high throughput binding assays, have shifted the theoretical focus from the prediction of new binding sites towards more quantitative models for the binding affinities of transcription factors, which can now be measured across whole genomes. Therefore we have developed a biophysical model which accounts for much of the observed variation in binding strength. Here we extend this framework to model not just the binding affinity, but also its distribution in various sequence backgrounds. This enables us to compare predicted affinities from different transcription factors, and to rank them according to their normalized affinity. What are the biological implications of such a ranking? We have demonstrated that many known associations between transcription factors and their respective targets appear as strong interactions. This provides a rationale to predict, for any given promoter region, those transcription factors which are most likely to be involved in its regulation.
DOI: 10.1016/0022-2836(87)90354-8
发表时间: 1987-02-20
影响因子: 5.6
作者:
BERG, OG;VONHIPPEL, PH
通讯作者: VONHIPPEL, PH
DOI: 10.1126/science.1075090
发表时间: 2002-10-25
期刊: SCIENCE
影响因子: 56.9
作者:
Lee, TI;Rinaldi, NJ;Young, RA
通讯作者: Young, RA
DOI: 10.1074/jbc.m304355200
发表时间: 2003-08-29
影响因子: 4.8
作者:
Hisamatsu, T;Suzuki, M;Podolsky, DK
通讯作者: Podolsky, DK
DOI: 10.1101/gad.8.13.1514
发表时间: 1994-07-01
影响因子: 10.5
作者:
JOHNSON, DG;OHTANI, K;NEVINS, JR
通讯作者: NEVINS, JR
DOI: 10.1126/science.290.5500.2306
发表时间: 2000-12-22
期刊: SCIENCE
影响因子: 56.9
作者:
Ren, B;Robert, F;Young, RA
通讯作者: Young, RA