IDENTIFYING PROTEIN-BINDING SITES FROM UNALIGNED DNA FRAGMENTS

IDENTIFYING PROTEIN-BINDING SITES FROM UNALIGNED DNA FRAGMENTS
复制标题

DOI:
10.1073/pnas.86.4.1183
复制
发表时间:
1989-02-01
影响因子:
11.1
通讯作者:
HARTZELL, GW
HARTZELL, GW
中科院分区:
综合性期刊1区
文献类型:
--
作者:
STORMO, GD;HARTZELL, GW

文献摘要

被引文献

相似文献

随着大规模测序项目的开展,仅从序列中确定DNA序列内的重要特征的能力变得至关重要。我们提出了一种方法,可以应用于识别的DNA结合蛋白的识别模式的问题,只给出了一个集合的测序的DNA片段,每个已知的包含在它的某个地方,该蛋白质的结合位点。不需要关于这些片段内的结合位点的位置或取向的信息。该方法比较大量可能的结合位点比对的“信息内容”,以得到结合位点模式的矩阵表示。蛋白质的特异性被表示为矩阵,而不是一个共识序列,允许模式,是典型的调节蛋白结合位点被确定。该方法的可靠性随着序列数量的增加而提高,但所需时间仅随序列数量线性增加。一个例子,使用已知的cAMP受体蛋白结合位点,说明了该方法。
The ability to determine important features within DNA sequences from the sequences alone is becoming essential as large-scale sequenching projects are being undertaken. We present a method that can be applied to the problem of identifying the recognition pattern for a DNA-binding protein given only a collection of sequenced DNA fragments, each known to contain somewhere within it a binding site for that protein. Information about the position or orientation of the binding sites within those fragments is not needed. The method compares the "information content" of a large number of possible binding site alignments to arrive at a matrix representation of the binding site pattern. The specificity of the protein is represented as a matrix, rather than a consenus sequence, allowing patterns that are typical of regulatory protein-binding sites to be identified. The reliability of the method improves as the number of sequences increases, but the time required increases only linearly with the number of sequences. An example, using known cAMP receptor protein-binding sites, illustrates the method.