Alignment of protein sequences by their profiles

Alignment of protein sequences by their profiles
复制标题

DOI:
10.1110/ps.03379804
复制
发表时间:
2004-04-01
期刊:
影响因子:
8
通讯作者:
Sali, A
Sali, A
中科院分区:
生物学3区
文献类型:
--
作者:
Marti-Renom, MA;Madhusudhan, MS;Sali, A

文献摘要

被引文献

相似文献

通过在比较中包括其他可检测到的相关序列,可以提高两个蛋白质序列之间比对的准确性。我们对这种方法进行优化和基准测试,该方法依赖于两个多重序列比对,每个序列比对包括两个蛋白质序列之一。 MODELLER 的 SALIGN 命令中实现了 13 种不同的协议,用于创建和比较与多个序列比对相对​​应的概况。 200 个基于结构的配对序列比对(序列同一性低于 40%)的测试集用于对 13 个协议以及许多先前描述的序列比对方法(包括通过 BLAST 进行的启发式配对序列比对)进行基准测试。通过全局动态规划进行成对序列比对,通过 MODELLER 的 ALIGN 命令使用仿射空位罚分函数,通过 PSI-BLAST 进行序列轮廓比对,在 SAM 和 LOBSTER 中实现的隐马尔可夫模型方法,通过 SEA 依赖于预测局部结构的成对序列比对,以及通过 CLUSTALW 和 COMPASS 进行多序列比对。最佳新协议的对齐精度明显优于其他测试方法。例如,相对于最佳方案基于结构的比对,正确比对的残基比例为 56%,这可以与其他方法的准确度分别为 26%、42%、43%、48%、50%、49%、43% 和 43% 进行比较。新方法目前应用于所有已知序列的大规模比较蛋白质结构建模。
The accuracy of an alignment between two protein sequences can be improved by including other detectably related sequences in the comparison. We optimize and benchmark such an approach that relies on aligning two multiple sequence alignments, each one including one of the two protein sequences. Thirteen different protocols for creating and comparing profiles corresponding to the multiple sequence alignments are implemented in the SALIGN command of MODELLER. A test set of 200 pairwise, structure-based alignments with sequence identities below 40% is used to benchmark the 13 protocols as well as a number of previously described sequence alignment methods, including heuristic pairwise sequence alignment by BLAST. pairwise sequence alignment by global dynamic programming with an affine gap penalty function by the ALIGN command of MODELLER, sequence-profile alignment by PSI-BLAST, Hidden Markov Model methods implemented in SAM and LOBSTER, pairwise sequence alignment relying on predicted local structure by SEA, and multiple sequence alignment by CLUSTALW and COMPASS. The alignment accuracies of the best new protocols were significantly better than those of the other tested methods. For example, the fraction of the correctly aligned residues relative to the structure-based alignment by the best protocol is 56%, which can be compared with the accuracies of 26%, 42%, 43%, 48%, 50%, 49%, 43%, and 43% for the other methods, respectively. The new method is currently applied to large-scale comparative protein structure modeling of all known sequences.