kClust: fast and sensitive clustering of large protein sequence databases.

kClust: fast and sensitive clustering of large protein sequence databases.
复制标题

DOI:
10.1186/1471-2105-14-248
复制
发表时间:
2013-08-15
期刊:
影响因子:
3
通讯作者:
Söding J
Söding J
中科院分区:
生物学4区
文献类型:
--
作者:
Hauser M;Mayer CE;Söding J

文献摘要

参考文献

被引文献

相似文献

在高通量测序的快速发展的推动下,公共序列数据库的规模每两年翻一番。搜索更大、更冗余的数据库变得越来越低效。聚类可以帮助将序列组织成同源和功能相似的组,并且可以提高同源性搜索的速度、灵敏度和可读性。然而,由于聚类时间是序列数量的二次函数,标准的序列搜索方法变得不切实际。在这里,我们提出了一种方法来聚类大型蛋白质序列数据库,如UniProt在几天内下降到20%-30%的最大成对序列同一性。kClust的速度和灵敏度归功于一个无重复的预过滤器,它计算序列对之间所有相似的6-mer的累积得分,以及一个对相似的4-mer进行操作的动态编程算法。为了进一步提高灵敏度,kClust可以在配置文件-序列比较模式下运行,配置文件是从上一次kClust迭代的聚类中计算出来的。kClust比基于NCBI BLAST的聚类快两到三个数量级,并且在20%-30%最大成对序列同一性的多域序列上,它实现了相当的灵敏度和较低的错误发现率。在错误发现率、灵敏度和速度方面,它也优于CD-HIT和UCLUST。kClust满足了对快速、灵敏和准确的工具的需求,可以将大型蛋白质序列数据库聚类到低于30%的序列同一性。kClust在GPL下免费提供,网址为http://toolkit.lmb.uni-muenchen.de/pub/kClust/。
Fueled by rapid progress in high-throughput sequencing, the size of public sequence databases doubles every two years. Searching the ever larger and more redundant databases is getting increasingly inefficient. Clustering can help to organize sequences into homologous and functionally similar groups and can improve the speed, sensitivity, and readability of homology searches. However, because the clustering time is quadratic in the number of sequences, standard sequence search methods are becoming impracticable. Here we present a method to cluster large protein sequence databases such as UniProt within days down to 20%–30% maximum pairwise sequence identity. kClust owes its speed and sensitivity to an alignment-free prefilter that calculates the cumulative score of all similar 6-mers between pairs of sequences, and to a dynamic programming algorithm that operates on pairs of similar 4-mers. To increase sensitivity further, kClust can run in profile-sequence comparison mode, with profiles computed from the clusters of a previous kClust iteration. kClust is two to three orders of magnitude faster than clustering based on NCBI BLAST, and on multidomain sequences of 20%–30% maximum pairwise sequence identity it achieves comparable sensitivity and a lower false discovery rate. It also compares favorably to CD-HIT and UCLUST in terms of false discovery rate, sensitivity, and speed. kClust fills the need for a fast, sensitive, and accurate tool to cluster large protein sequence databases to below 30% sequence identity. kClust is freely available under GPL at http://toolkit.lmb.uni-muenchen.de/pub/kClust/.
DOI: 10.1093/nar/gkm107
发表时间: 2007
影响因子: 14.9
作者:
Przybylski D;Rost B
通讯作者: Rost B
DOI: 10.1038/nature08821
发表时间: 2010-03-04
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1093/bioinformatics/bts565
发表时间: 2012-12-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Fu L;Niu B;Zhu Z;Wu S;Li W
通讯作者: Li W
DOI: 10.1093/bioinformatics/btr447
发表时间: 2011-09-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Bao, Ergude;Jiang, Tao;Girke, Thomas
通讯作者: Girke, Thomas
DOI: 10.1093/nar/28.1.270
发表时间: 2000-01-01
影响因子: 14.9
作者:
Krause, A;Stoye, J;Vingron, M
通讯作者: Vingron, M