Swarm v2: highly-scalable and high-resolution amplicon clustering.

Swarm v2: highly-scalable and high-resolution amplicon clustering.
复制标题

DOI:
10.7717/peerj.1420
复制
发表时间:
2015
期刊:
影响因子:
2.7
通讯作者:
Dunthorn M
Dunthorn M
中科院分区:
生物学3区
文献类型:
--
作者:
Mahé F;Rognes T;Quince C;de Vargas C;Dunthorn M

文献摘要

被引文献

相似文献

在此之前,我们提出了一种新颖的开源扩增子聚类程序Sarm v1,它产生了精细的分子操作分类单元(OTU),没有任意的全局聚类阈值和输入顺序依赖。SWARM v1的初始阶段使用具有局部聚类阈值(D)的迭代单链,随后的阶段使用集群的内部丰度结构来打破连锁的OTU。在这里,我们提出了Sarm v2,它有两个重要的新特征:(1)d=1的新算法,它允许程序的计算时间随着数据量的增加而线性扩展;(2)新的挑剔选项,通过将低丰富的OTU(例如,单线程和双线程)嫁接到较大的OTU上来减少分组。Sarm v2还直接整合了聚集和破裂阶段,取消了d=0的测序读数,以FASTA格式输出OTU代表,并将单个OTU绘制为二维网络。
Previously we presented Swarm v1, a novel and open source amplicon clustering program that produced fine-scale molecular operational taxonomic units (OTUs), free of arbitrary global clustering thresholds and input-order dependency. Swarm v1 worked with an initial phase that used iterative single-linkage with a local clustering threshold (d), followed by a phase that used the internal abundance structures of clusters to break chained OTUs. Here we present Swarm v2, which has two important novel features: (1) a new algorithm for d = 1 that allows the computation time of the program to scale linearly with increasing amounts of data; and (2) the new fastidious option that reduces under-grouping by grafting low abundant OTUs (e.g., singletons and doubletons) onto larger ones. Swarm v2 also directly integrates the clustering and breaking phases, dereplicates sequencing reads with d = 0, outputs OTU representatives in fasta format, and plots individual OTUs as two-dimensional networks.