Sensitive sequence comparison as protein function predictor.

Sensitive sequence comparison as protein function predictor.
复制标题

DOI:
10.1142/9789814447331_0005
复制
发表时间:
1999-12
影响因子:
--
通讯作者:
K. Pawłowski;L. Jaroszewski;L. Rychlewski;A. Godzik
K. Pawłowski;L. Jaroszewski;L. Rychlewski;A. Godzik
中科院分区:
--
文献类型:
--
作者:
K. Pawłowski;L. Jaroszewski;L. Rychlewski;A. Godzik

文献摘要

相似文献

基于通过高序列相似性识别的假定同源性的蛋白质功能分配通常用于基因组分析。序列比较算法灵敏度的改进已经达到了这样的程度:具有以前无法检测到的序列相似性的蛋白质,例如 10-15% 的相同残基,有时可以被归类为相似的。这些蛋白质之间有什么关系?它们有可能是同源的吗?检测这种相似性有什么实际意义呢?本文针对大肠杆菌基因组中已充分表征的蛋白质,对序列相似性和功能相似性之间的关系进行了简化分析。使用基于酶的 E.C. 分类的功能相似性的简单测量,结果表明它与通过比对得分的统计显着性测量的序列相似性良好相关。按照这个标准,相似的蛋白质,即使在序列同一性较低的情况下,也比随机选择的蛋白质对有更大的机会具有相似的功能。讨论了这些规则的有趣例外。
Protein function assignments based on postulated homology as recognized by high sequence similarity are used routinely in genome analysis. Improvements in sensitivity of sequence comparison algorithms got to the point, that proteins with previously undetectable sequence similarity, such as for instance 10-15% of identical residues, sometimes can be classified as similar. What is the relation between such proteins? Is it possible that they are homologous? What is the practical significance of detecting such similarities? A simplified analysis of the relation between sequence similarity and function similarity is presented here for the well-characterized proteins from the E. coli genome. Using a simple measure of functional similarity based on E.C. classification of enzymes, it is shown that it correlates well with sequence similarity measured by statistical significance of the alignment score. Proteins, similar by this standard, even in cases of low sequence identity, have a much larger chance of having similar function than the randomly chosen protein pairs. Interesting exceptions to these rules are discussed.