SEARCHING PROTEIN-SEQUENCE LIBRARIES - COMPARISON OF THE SENSITIVITY AND SELECTIVITY OF THE SMITH-WATERMAN AND FASTA ALGORITHMS

SEARCHING PROTEIN-SEQUENCE LIBRARIES - COMPARISON OF THE SENSITIVITY AND SELECTIVITY OF THE SMITH-WATERMAN AND FASTA ALGORITHMS
复制标题

DOI:
10.1016/0888-7543(91)90071-l
复制
发表时间:
1991-11-01
期刊:
影响因子:
4.4
通讯作者:
PEARSON, WR
PEARSON, WR
中科院分区:
生物学3区
文献类型:
--
作者:
PEARSON, WR

文献摘要

被引文献

相似文献

使用国家生物医学研究基金会/蛋白质鉴定资源(PIR)蛋白质序列数据库中提供的超家族分类评价FASTA和Smith-Waterman蛋白质序列比较算法的灵敏度和选择性。来自PIR数据库中具有20个或更多成员的34个超家族中的每一个的序列与蛋白质序列数据库进行比较。使用FASTA程序或Smith-Waterman局部相似性算法确定相关和不相关序列的相似性得分。这两组相似性得分用于评估两种比较算法识别远缘相关蛋白质序列的能力。FASTA程序使用kS = 2的灵敏度设置,以及史密斯-沃特曼算法的34个超家族中的19个。通过设置kmax = 1来增加灵敏度,使得FASTA在另外7个超家族上的表现与Smith-Waterman一样好。严格的Smith-Waterman方法在包括球蛋白、免疫球蛋白可变区、钙调素和质体蓝素在内的8个超家族上的表现优于Kmax = 1的FASTA方法。研究了提高FASTA灵敏度的几种策略。灵敏度的最大改善是通过优化每个文库序列发现的最佳初始区域周围的条带来实现的。对于除了球蛋白和免疫球蛋白可变区之外的每个超家族,这种策略与完整的史密斯-沃特曼一样敏感。对于一些序列,通过在用于鉴定初始区域的查找表中包括保守但不相同的残基来实现额外的灵敏度。
The sensitivity and selectivity of the FASTA and the Smith-Waterman protein sequence comparison algorithms were evaluated using the superfamily classification provided in the National Biomedical Research Foundation/Protein Identification Resource (PIR) protein sequence database. Sequences from each of the 34 superfamilies in the PIR database with 20 or more members were compared against the protein sequence database. The similarity scores of the related and unrelated sequences were determined using either the FASTA program or the Smith-Waterman local similarity algorithm. These two sets of similarity scores were used to evaluate the ability of the two comparison algorithms to identify distantly related protein sequences. The FASTA program using thektup = 2sensitivity setting performed as well as the Smith-Waterman algorithm for 19 of the 34 superfamilies. Increasing the sensitivity by settingktup = 1allowed FASTA to perform as well as Smith-Waterman on an additional 7 superfamilies. The rigorous Smith-Waterman method performed better than FASTA withktup = 1on 8 superfamilies, including the globins, immunoglobulin variable regions, calmodulins, and plastocyanins. Several strategies for improving the sensitivity of FASTA were examined. The greatest improvement in sensitivity was achieved by optimizing a band around the best initial region found for every library sequence. For every superfamily except the globins and immunoglobulin variable regions, this strategy was as sensitive as a full Smith-Waterman. For some sequences, additional sensitivity was achieved by including conserved but nonidentical residues in the lookup table used to identify the initial region.