Combining sensitive database searches with multiple intermediates to detect distant homologues

Combining sensitive database searches with multiple intermediates to detect distant homologues
复制标题

DOI:
10.1093/protein/12.2.95
复制
发表时间:
1999-02-01
期刊:
PROTEIN ENGINEERING
影响因子:
--
通讯作者:
Swindells, MB
Swindells, MB
中科院分区:
其他
文献类型:
--
作者:
Salamov, AA;Suwa, M;Swindells, MB

文献摘要

被引文献

相似文献

使用Cath结构分类的数据,我们评估了BLASTP、FASTA、Smith-Waterman和Gap-BLAST算法,开发了一种可移植的归一化方案,并确定了数据库搜索的安全阈值。在评估的四种方法中,FASTA、Smith-Waterman和Gap-BLAST的表现相似,而BLAST的敏感度要低得多。中间序列搜索的引入大大改善了结果。当对blastp无法识别的一组关系进行测试时,中间序列能够找到仅由Smith-Waterman算法识别的关系数量的两倍。然而,我们发现,使用中间体的好处在每个家族之间有很大的差异,不仅取决于可用序列的数量,而且还取决于它们的多样性。为了进一步提高灵敏度,开发了多中间序列搜索(MISTH)程序。当对来自广泛的同源家族的1906个病例进行评估时,MISTH能够识别出241个额外的关系。未命中使用全范围的序列多样性来检测额外的关系,但不考虑任何结构特定的信息。因此,它比折叠识别和穿线方法更适用,后者需要已知结构的库。
Using data from the CATH structure classification, we have assessed the blastp, fasta, smith-waterman and gapped-blast algorithms, developed a portable normalization scheme and identified safe thresholds for database searching. Of the four methods assessed, fasta, smith-waterman and gapped-blast perform similarly, whereas the sensitivity of blastp was much lower. Introduction of an intermediate sequence search substantially improved the results. When tested on a set of relationships that could not be identified by blastp, intermediate sequences were able to find double the number of relationships identified by the smith-waterman algorithm alone. However, we found that the benefit of using intermediates varied considerably between each family and depended not only on the number of available sequences, but also their diversity. In an attempt to increase sensitivity further, a multiple intermediate sequence search (MISS) procedure was developed. When assessed on 1906 cases from a wide range of homologous families that could not be detected by the previous approaches, MISS was able to identify 241 additional relationships. MISS uses the full extent of sequence diversity to detect additional relationships, but does not consider any structure-specific information. For this reason, it is more generally applicable than fold recognition and threading methods, which require a library of known structures.