CD-HIT: accelerated for clustering the next-generation sequencing data.
CD-HIT: accelerated for clustering the next-generation sequencing data.
复制标题
DOI:
10.1093/bioinformatics/bts565
复制
发表时间:
2012-12-01
期刊:
影响因子:
--
通讯作者:
Li W
中科院分区:
文献类型:
--
作者:
Fu L;Niu B;Zhu Z;Wu S;Li W
Summary: CD-HIT is a widely used program for clustering biological sequences to reduce sequence redundancy and improve the performance of other sequence analyses. In response to the rapid increase in the amount of sequencing data produced by the next-generation sequencing technologies, we have developed a new CD-HIT program accelerated with a novel parallelization strategy and some other techniques to allow efficient clustering of such datasets. Our tests demonstrated very good speedup derived from the parallelization for up to ∼24 cores and a quasi-linear speedup for up to ∼8 cores. The enhanced CD-HIT is capable of handling very large datasets in much shorter time than previous versions. Availability: http://cd-hit.org. Contact: liwz@sdsc.edu Supplementary information: Supplementary data are available at Bioinformatics online.
登录
查看更多内容
影响因子:
64.8
作者:
通讯作者:
--
影响因子:
5.8
作者:
Li, WZ;Jaroszewski, L;Godzik, A
通讯作者:
Godzik, A
影响因子:
14.9
作者:
Sun S;Chen J;Li W;Altintas I;Lin A;Peltier S;Stocks K;Allen EE;Ellisman M;Grethe J;Wooley J
通讯作者:
Wooley J
影响因子:
3
作者:
Niu B;Fu L;Sun S;Li W
通讯作者:
Li W
影响因子:
5.8
作者:
Edgar, Robert C.
通讯作者:
Edgar, Robert C.