Clustering of highly homologous sequences to reduce the size of large protein databases

Clustering of highly homologous sequences to reduce the size of large protein databases
复制标题

DOI:
10.1093/bioinformatics/17.3.282
复制
发表时间:
2001-03-01
期刊:
影响因子:
5.8
通讯作者:
Godzik, A
Godzik, A
中科院分区:
生物学3区
文献类型:
--
作者:
Li, WZ;Jaroszewski, L;Godzik, A

文献摘要

被引文献

相似文献

我们提出了一个快速灵活的程序,用于在不同的序列身份水平下聚类大蛋白质数据库。在高端PC上,超过560 000个序列的非冗余蛋白数据库的全序列比较和聚类的全序序列比较和聚类所需少于2小时。输出数据库(包括代表性序列)可用于更有效,更敏感的数据库搜索。
We present a fast and flexible program for clustering large protein databases at different sequence identity levels. It takes less than 2 h for the all-against-all sequence comparison and clustering of the non-redundant protein database of over 560 000 sequences on a high-end PC. The output database, including only the representative sequences, can be used for more efficient and sensitive database searches.