SELECTION OF DNA-BINDING SITES BY REGULATORY PROTEINS - STATISTICAL-MECHANICAL THEORY AND APPLICATION TO OPERATORS AND PROMOTERS

SELECTION OF DNA-BINDING SITES BY REGULATORY PROTEINS - STATISTICAL-MECHANICAL THEORY AND APPLICATION TO OPERATORS AND PROMOTERS
复制标题

DOI:
10.1016/0022-2836(87)90354-8
复制
发表时间:
1987-02-20
影响因子:
5.6
通讯作者:
VONHIPPEL, PH
VONHIPPEL, PH
中科院分区:
生物学2区
文献类型:
--
作者:
BERG, OG;VONHIPPEL, PH

文献摘要

被引文献

相似文献

我们提出了一种统计机械选择理论,用于一组特定DNA调控位点的序列分析,这使得预测该位点中单个碱基对选择与特定活性(亲和力)之间的关系成为可能。该理论是基于这样的假设:特定的DNA序列被选择出来,以符合蛋白质结合(或活性)的某些要求,并且所有能够满足这一要求的序列都同样有可能出现。在大多数情况下,已知特定DNA结合蛋白的特定DNA序列的数量非常少,我们将详细讨论这导致的小样本不确定性。当应用于噬菌体lambda中交叉抑制因子的结合位点时,该理论仅从序列统计量就能预测它们的秩序结合亲和力,与实测值基本一致。然而,这样一个小样本(只有6个已知地点)产生的统计不确定性限制了结果的数量级比较。当应用于更大的大肠杆菌启动子序列样本时,该理论预测了Mulligan et al.(1984)观察到的体外活性(k2KB值)与同源性评分(接近一致序列)之间的相关性。对启动子样本中碱基对频率的分析与假设一致,即位点上不同位置的碱基对对特定活性的独立贡献,除了在讨论的少数边缘情况下。当启动子位点根据预测活性排序时,它们似乎符合高斯分布,这是在提供一定平均活性的约束下要求最大序列可变性的结果。该理论允许我们将具有特定活性的特定位点的数量与基因组中随机出现的预期数量进行比较。虽然强启动子被“过度指定”,即它们随机出现的概率非常低,但具有弱启动子性质的随机序列预计会大量出现。由此得出的结论是,除了初级序列识别外,功能特异性还基于其他特性;讨论了一些可能性。最后,我们证明了Schneider等人(1986)定义的序列信息可以直接使用(至少在平衡结合位点的情况下)来估计基因组中随机“假位点”特异性结合的蛋白质分子的数量。这提供了碱基对序列统计与von Hippel和Berg(1986)定义的体内功能特异性之间的联系。
We present a statistical-mechanical selection theory for the sequence analysis of a set of specific DNA regulatory sites that makes it possible to predict the relationship between individual base-pair choices in the site and specific activity (affinity). The theory is based on the assumption that specific DNA sequences have been selected to conform to some requirement for protein binding (or activity), and that all sequences that can fulfill this requirement are equally likely to occur. In most cases, the number of specific DNA sequences that are known for a certain DNA-binding protein is very small, and we discuss in detail the small sample uncertainties that this leads to. When applied to the binding sites for cro repressor in phage lambda, the theory can predict, from the sequence statistics alone, their rank order binding affinities in reasonable agreement with measured values. However, the statistical uncertainty generated by such a small sample (only 6 sites known) limits the result to order-of-magnitude comparisons. When applied to the much larger sample of Escherichia coli promoter sequences, the theory predicts the correlation between in vitro activity (k2KB values) and homology score (closeness to the consensus sequence) observed by Mulligan et al. (1984). The analysis of base-pair frequencies in the promoter sample is consistent with the assumption that base-pairs at different positions in the sites contribute independently to the specific activity, except in a few marginal cases that are discussed. When the promoter sites are ordered according to predicted activities, they seem to conform to the Gaussian distribution that results from a requirement for maximal sequences variability within the constraint of providing a certain average activity. The theory allows us to compare the number of specific sites with a certain activity to the number that would be expected from random occurrence in the genome. While strong promoters are "overspecified", in the sense that their probability of random occurrence is very low, random sequences with weak promoter-like properties are expected to occur in very large numbers. This leads to the conclusion that functional specificity is based on other properties in additional to primary sequence recognition; some possibilities are discussed. Finally, we show that the sequence information, as defined by Schneider et al. (1986), can be used directly (at least in the case of equilibrium binding sites) to estimate the number of protein molecules that are specifically bound at random "pseudosites" in the genome. This provides the connection between base-pair sequences statistics and functional in vivo specificity as defined by von Hippel and Berg (1986).