Empirical statistical estimates for sequence similarity searches

Empirical statistical estimates for sequence similarity searches
复制标题

DOI:
10.1006/jmbi.1997.1525
复制
发表时间:
1998-02-13
影响因子:
5.6
通讯作者:
Pearson, WR
Pearson, WR
中科院分区:
生物学2区
文献类型:
--
作者:
Pearson, WR

文献摘要

被引文献

相似文献

对FASTA序列比较程序包进行了修改,以便为有空位的局部序列相似性分数提供准确的统计估计。这些估计是在对文库序列长度的预期效果进行校正后,从不相关序列的局部相似性分数的平均值和方差的极值分布得到的。这种方法允许对蛋白质/蛋白质、DNA/DNA和蛋白质/翻译DNA比较的FASTA和Smith-Waterman相似性分数进行准确的估计。使用Fasta和Smith-Waterman分数对54个蛋白质家族的统计估计的准确性进行了总结。根据相似性得分的分布计算的概率估计通常是保守的,使用Altschul-Gish lambda、K和H参数计算的概率也是保守的。使用来自PIR39数据库的54个蛋白质超家族和来自ProSite/SwissProt Rel的110个蛋白质家族评估了几种校正文库序列长度相似性分数的替代方法的性能。34个数据库。回归评分和Altschul-Gish评分的表现都明显好于未封存的Smith-Waterman或Fasta相似性评分。当使用ProSite/SwissProt测试集时,回归尺度的分数表现略好;当使用PIR数据库时,Altschul-Gish尺度的分数表现最好。因此,长度校正的相似性分数提高了数据库搜索的敏感度。从数据库搜索中通常遇到的数千个不相关序列的相似性分数的分布导出的统计参数提供了可用于推断序列同源性的统计意义的准确估计。(C)1998年学术出版社有限公司。
The FASTA package of sequence comparison programs has been modified to provide accurate statistical estimates for local sequence similarity scores with gaps. These estimates are derived using the extreme value distribution from the mean and variance of the local similarity scores of unrelated sequences after the scores have been corrected for the expected effect of library sequence length. This approach allows accurate estimates to be calculated for both FASTA and Smith-Waterman similarity scores for protein/protein, DNA/DNA, and protein/translated-DNA comparisons. The accuracy of the statistical estimates is summarized for 54 protein families using FASTA and Smith-Waterman scores. Probability estimates calculated from the distribution of similarity scores are generally conservative, as are probabilities calculated using the Altschul-Gish lambda, K, and H parameters. The performance of several alternative methods for correcting similarity scores for library-sequence length was evaluated using 54 protein superfamilies from the PIR39 database and 110 protein families from the Prosite/SwissProt rel. 34 database. Both regression-scaled and Altschul-Gish scaled scores perform significantly better than unsealed Smith-Waterman or FASTA similarity scores. When the Prosite/SwissProt test set is used, regression-scaled scores perform slightly better; when the PIR database is used, Altschul-Gish scaled scores perform best. Thus, length-corrected similarity scores improve the sensitivity of database searches. Statistical parameters that are derived from the distribution of similarity scores from the thousands of unrelated sequences typically encountered in a database search provide accurate estimates of statistical significance that can be used to infer sequence homology. (C) 1998 Academic Press Limited.