AN EXPECTATION MAXIMIZATION (EM) ALGORITHM FOR THE IDENTIFICATION AND CHARACTERIZATION OF COMMON SITES IN UNALIGNED BIOPOLYMER SEQUENCES

AN EXPECTATION MAXIMIZATION (EM) ALGORITHM FOR THE IDENTIFICATION AND CHARACTERIZATION OF COMMON SITES IN UNALIGNED BIOPOLYMER SEQUENCES
复制标题

DOI:
10.1002/prot.340070105
复制
发表时间:
1990-01-01
影响因子:
2.9
通讯作者:
REILLY, AA
REILLY, AA
中科院分区:
生物学4区
文献类型:
--
作者:
LAWRENCE, CE;REILLY, AA

文献摘要

被引文献

相似文献

统计方法的识别和表征的蛋白质结合位点的一组未对齐的DNA片段。每个序列必须包含至少一个共同位点。不需要对网站进行调整。相反,在网站的位置的不确定性处理采用缺失信息的原则,开发一个“期望最大化”(EM)算法。这种方法允许同时识别位点和表征结合基序。该算法的可靠性随着片段的数量而增加,但计算量仅线性增加。该方法用一个例子说明,使用已知的环磷酸腺苷受体蛋白(CRP)结合位点。最后的基序用于搜索未发现的CRP结合位点。
Statistical methodology for the identification and characterization of protein binding sites in a set of unaligned DNA fragments is presented. Each sequence must contain at least one common site. No alignment of the sites is required. Instead, the uncertainty in the location of the sites is handled by employing the missing information principle to develop an "expectation maximization" (EM) algorithm. This approach allows for the simultaneous identification of the sites and characterization of the binding motifs. The reliability of the algorithm increases with the number of fragments, but the computations increase only linearly. The method is illustrated with an example, using known cyclic adenosine monophosphate receptor protein (CRP) binding sites. The final motif is utilized in a search for undiscovered CRP binding sites.