Identification of the binding sites of regulatory proteins in bacterial genomes

Identification of the binding sites of regulatory proteins in bacterial genomes
复制标题

DOI:
10.1073/pnas.112341999
复制
发表时间:
2002-09-03
影响因子:
11.1
通讯作者:
Siggia, ED
Siggia, ED
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Li, H;Rhodius, V;Siggia, ED

文献摘要

被引文献

相似文献

我们提出了一种算法,可以从基因组的调控区域中提取许多不同转录因子的结合位点(由位置特异性权重矩阵表示),而不需要描绘共同调控的基因组。该算法利用了这样一个事实:细菌中的许多 DNA 结合蛋白与二分基序结合,其中有两个比中间区域更保守的短片段。它识别 W1NxW2 形式的所有统计显着模式,其中 W-1 和 W-2 是由 x 个任意碱基分隔的两个短寡核苷酸,并将它们分组为相似模式的簇。然后使用这些簇来导出假定的调节蛋白的定量识别图谱。对于给定的簇,该算法找到匹配序列以及基因组中的侧翼区域,并执行多序列比对以导出特定位置的权重矩阵。我们使用该算法分析了大肠杆菌基因组,发现了大约 1,500 个显着模式,从而产生了大约 160 个不同的特定位置权重矩阵。这些矩阵的一小部分与大约 60 个特征转录因子中三分之一的结合位点相匹配,具有很高的统计显着性。许多剩余的矩阵可能描述了未表征的转录因子的结合位点和调节子。这些矩阵的重要性是通过它们的特异性、预测位点的位置以及相应调节子的生物学功能来评估的,使我们能够提出假定的调节功能。该算法对于分析新测序的细菌基因组非常有效,而对于转录调控知之甚少。
We present an algorithm that extracts the binding sites (represented by position-specific weight matrices) for many different transcription factors from the regulatory regions of a genome, without the need for delineating groups of coregulated genes. The algorithm uses the fact that many DNA-binding proteins in bacteria bind to a bipartite motif with two short segments more conserved than the intervening region. It identifies all statistically significant patterns of the form W1NxW2, where W-1 and W-2 are two short oligonuclecitides separated by x arbitrary bases, and groups them into clusters of similar patterns. These clusters are then used to derive quantitative recognition profiles of putative regulatory proteins. For a given cluster, the algorithm finds the matching sequences plus the flanking regions in the genome and performs a multiple sequence alignment to derive position-specific weight matrices. We have analyzed the Escherichia coli genome with this algorithm and found approximate to1,500 significant patterns, which give rise to approximate to160 distinct position-specific weight matrices. A fraction of these matrices match the binding sites of one-third of the approximate to60 characterized transcription factors with high statistical significance. Many of the remaining matrices are likely to describe binding sites and regulons of uncharacterized transcription factors. The significance of these matrices was evaluated by their specificity, the location of the predicted sites, and the biological functions of the corresponding regulons, allowing us to suggest putative regulatory functions. The algorithm is efficient for analyzing newly sequenced bacterial genomes for which little is known about transcriptional regulation.