COMPARISON OF METHODS FOR SEARCHING PROTEIN-SEQUENCE DATABASES

COMPARISON OF METHODS FOR SEARCHING PROTEIN-SEQUENCE DATABASES
复制标题

DOI:
10.1002/pro.5560040613
复制
发表时间:
1995-06-01
期刊:
影响因子:
8
通讯作者:
PEARSON, WR
PEARSON, WR
中科院分区:
生物学3区
文献类型:
--
作者:
PEARSON, WR

文献摘要

被引文献

相似文献

我们比较了常用的序列比较算法,评分矩阵,和缺口罚分使用的方法,确定统计学上显着的性能差异。使用Smith-Waterman算法或FASTA的搜索灵敏度通过使用现代评分矩阵(例如BL 0 SUM 45 -55)和优化的空位罚分而不是常规的PAM 250矩阵而显著提高。通过用文库序列长度的对数来缩放相似性分数(ln()-scaling),可以获得更显著的改善。在最佳现代评分矩阵(BL 0 SUM 55或J 093)和最佳空位罚分(对于差距中的第一个残基为-12,对于另外的残基为-2)的情况下,Smith-Waterman和FASTA的表现显著优于BLASTP。使用ln()-缩放和最佳评分矩阵(BL 0 SUM 45或Gonnet 92)和空位罚分(-12,-1),严格的Smith-Waterman算法比BLASTP和FASTA表现更好,尽管使用Gonnet 92矩阵与FASTA的差异不显著。Ln()-缩放比基于库序列长度的其他简单函数的标准化表现更好。Ln()-标度也比基于归一化方差的评分更好,但对于BL 0 SUM 50和Gonnet 92矩阵,差异无统计学显著性。最佳评分矩阵和缺口罚分报告史密斯-沃特曼和FASTA,使用传统的或ln()-缩放的相似性得分。没有缺口延伸罚分、或没有缺口开放罚分、或缺口无限罚分的方法比最好的方法表现得明显更差。当使用部分查询序列时,PASTA和Smith-Waterman之间的性能差异并不显著。然而,最好的性能与完整的查询序列获得了史密斯-沃特曼算法和ln()-缩放。
We have compared commonly used sequence comparison algorithms, scoring matrices, and gap penalties using a method that identifies statistically significant differences in performance. Search sensitivity with either the Smith-Waterman algorithm or FASTA is significantly improved by using modern scoring matrices, such as BLOSUM45-55, and optimized gap penalties instead of the conventional PAM250 matrix. More dramatic improvement can be obtained by scaling similarity scores by the logarithm of the length of the library sequence (ln()-scaling). With the best modern scoring matrix (BLOSUM55 or JO93) and optimal gap penalties (-12 for the first residue in the gap and -2 for additional residues), Smith-Waterman and FASTA performed significantly better than BLASTP. With ln()-scaling and optimal scoring matrices (BLOSUM45 or Gonnet92) and gap penalties (-12, -1), the rigorous Smith-Waterman algorithm performs better than either BLASTP and FASTA, although with the Gonnet92 matrix the difference with FASTA was not significant. Ln()-scaling performed better than normalization based on other simple functions of library sequence length. Ln()-scaling also performed better than scores based on normalized variance, but the differences were not statistically significant for the BLOSUM50 and Gonnet92 matrices. Optimal scoring matrices and gap penalties are reported for Smith-Waterman and FASTA, using conventional or ln()-scaled similarity scores. Searches with no penalty for gap extension, or no penalty for gap opening, or an infinite penalty for gaps performed significantly worse than the best methods. Differences in performance between PASTA and Smith-Waterman were not significant when partial query sequences were used. However, the best performance with complete query sequences was obtained with the Smith-Waterman algorithm and ln()-scaling.