Search and clustering orders of magnitude faster than BLAST

Search and clustering orders of magnitude faster than BLAST
复制标题

DOI:
10.1093/bioinformatics/btq461
复制
发表时间:
2010-10-01
期刊:
影响因子:
5.8
通讯作者:
Edgar, Robert C.
Edgar, Robert C.
中科院分区:
生物学3区
文献类型:
--
作者:
Edgar, Robert C.

文献摘要

被引文献

相似文献

动机:生物序列数据正在迅速积累,促进了高通量序列分类方法的发展。结果:UBLAST和USEARCH是一种新的算法,能够以极高的速度对大型序列数据库进行敏感的局部和全局搜索。在实际应用中,它们通常比BLAST快几个数量级,尽管对远距离蛋白质关系的敏感性较低。UCLUST是一种新的聚类方法,它利用USEARCH将序列分配给聚类。与广泛使用的程序CD-HIT相比,UCLUST具有几个优势,包括更高的速度、更低的内存使用、更高的灵敏度、更低身份的聚类和更大数据集的分类。
Motivation: Biological sequence data is accumulating rapidly, motivating the development of improved high-throughput methods for sequence classification.Results: UBLAST and USEARCH are new algorithms enabling sensitive local and global search of large sequence databases at exceptionally high speeds. They are often orders of magnitude faster than BLAST in practical applications, though sensitivity to distant protein relationships is lower. UCLUST is a new clustering method that exploits USEARCH to assign sequences to clusters. UCLUST offers several advantages over the widely used program CD-HIT, including higher speed, lower memory use, improved sensitivity, clustering at lower identities and classification of much larger datasets.