Sequence clustering strategies improve remote homology recognitions while reducing search times

Sequence clustering strategies improve remote homology recognitions while reducing search times
复制标题

DOI:
10.1093/protein/15.8.643
复制
发表时间:
2002-08-01
期刊:
PROTEIN ENGINEERING
影响因子:
--
通讯作者:
Godzik, A
Godzik, A
中科院分区:
其他
文献类型:
--
作者:
Li, WZ;Jaroszewski, L;Godzik, A

文献摘要

被引文献

相似文献

序列数据库正在迅速增长,从而增加了蛋白质序列空间的覆盖范围,但这种覆盖范围是不均匀的,因为大多数测序工作集中在少数生物体上。由此产生的序列空间的粒度为基于配置文件的序列比较程序带来了许多问题。在本文中,我们提出了几种策略,解决这些问题,并在同一时间加快同源蛋白质的搜索和提高能力的配置文件的方法来识别遥远的同源性。我们的策略之一,结合数据库聚类,删除高度冗余的序列,和两步PSI-BLAST(PDB-BLAST),分离序列空间的配置文件组成和空间的同源性搜索。这些策略的组合将远缘同源性排除提高了100%以上,而仅使用标准PSI-BLAST搜索的10%的CPU时间。另一种方法,中间谱搜索,允许探索额外的搜索方向,通常是由非常不同的家庭内的大蛋白质亚家族占主导地位。所有方法都使用大型折叠识别基准进行评估。
Sequence databases are rapidly growing, thereby increasing the coverage of protein sequence space, but this coverage is uneven because most sequencing efforts have concentrated on a small number of organisms. The resulting granularity of sequence space creates many problems for profile-based sequence comparison programs. In this paper, we suggest several strategies that address these problems, and at the same time speed up the searches for homologous proteins and improve the ability of profile methods to recognize distant homologies. One of our strategies combines database clustering, which removes highly redundant sequence, and a two-step PSI-BLAST (PDB-BLAST), which separates sequence spaces of profile composition and space of homology searching. The combination of these strategies improves distant homology recognitions by more than 100%, while using only 10% of the CPU time of the standard PSI-BLAST search. Another method, intermediate profile searches, allows for the exploration of additional search directions that are normally dominated by large protein sub-families within very diverse families. All methods are evaluated with a large fold-recognition benchmark.