Comparison of sequence profiles. Strategies for structural predictions using sequence information

Comparison of sequence profiles. Strategies for structural predictions using sequence information
复制标题

DOI:
10.1110/ps.9.2.232
复制
发表时间:
2000-02-01
期刊:
影响因子:
8
通讯作者:
Godzik, A
Godzik, A
中科院分区:
生物学3区
文献类型:
--
作者:
Rychlewski, L;Jaroszewski, L;Godzik, A

文献摘要

被引文献

相似文献

蛋白质之间的远距离同源性通常只有在解决了两种蛋白质的三维结构之后才能被发现。这类蛋白质的序列差异可能如此之大,以至于简单地比较它们的序列无法识别任何相似之处。新一代敏感的比对工具使用整个同源家族(图谱)的平均序列来检测这种同源性。在结构相似、序列相似性很小的大集合蛋白质上,比较了几种算法,包括最新一代的BLAST算法和我们团队中使用的BASIC算法,该算法用于分配来自几个基因组的蛋白质的折叠预测。根据相似性程度对基准中的蛋白质进行分类,这使得我们可以证明,新算法的大部分改进是针对功能相似性较强的蛋白质而实现的,在识别远距离折叠相似性方面几乎没有进展,同时还表明轮廓计算的细节对其识别远距离同源序列的敏感性有很大影响。最重要的选择是如何包含来自不同家族成员的信息,避免产生错误预测,同时考虑到一个家族内的整个序列差异。PSI-BLAST采取了保守的方法,从家庭的核心成员那里获得了一个简介,提供了坚实的改进,几乎没有任何错误的预测。Basic通过增加不同家庭成员的权重来努力提高敏感度,并以较低的可靠性为代价。这里介绍的一种新的FFAS算法使用了一种新的轮廓生成程序,该程序考虑了家族内的所有关系,并将基本灵敏度与PSI-BLAST类似的可靠性进行了匹配。
Distant homologies between proteins are often discovered only after three-dimensional structures of both proteins are solved. The sequence divergence for such proteins can be so large that simple comparison of their sequences fails to identify any similarity. New generation of sensitive alignment tools use averaged sequences of entire homologous families (profiles) to detect such homologies. Several algorithms, including the newest generation of BLAST algorithms and BASIC, an algorithm used in our group to assign fold predictions for proteins from several genomes, are compared to each other on the large set of structurally similar proteins with little sequence similarity. Proteins in the benchmark are classified according to the level of their similarity, which allows us to demonstrate that most of the improvement of the new algorithms is achieved for proteins with strong functional similarities, with almost no progress in recognizing distant fold similarities.It is also shown that details of profile calculation strongly influence its sensitivity in recognizing distant homologies. The most important choice is how to include information from diverging members of the family, avoiding generating false predictions, while accounting for entire sequence divergence within a family. PSI-BLAST takes a conservative approach, deriving a profile from core members of the family, providing a solid improvement without almost any false predictions. BASIC strives for better sensitivity by increasing the weight of divergent family members and paying the price in lower reliability. A new FFAS algorithm introduced here uses a new procedure for profile generation that takes into account all the relations within the family and matches BASIC sensitivity with PSI-BLAST like reliability.