Toward accurate molecular identification of species in complex environmental samples: testing the performance of sequence filtering and clustering methods.

Toward accurate molecular identification of species in complex environmental samples: testing the performance of sequence filtering and clustering methods.
复制标题

DOI:
10.1002/ece3.1497
复制
发表时间:
2015-06
影响因子:
2.6
通讯作者:
Cristescu, Melania E.
Cristescu, Melania E.
中科院分区:
生物学2区
文献类型:
--
作者:
Flynn, Jullien M.;Brown, Emily A.;Chain, Frederic J. J.;MacIsaac, Hugh J.;Cristescu, Melania E.

文献摘要

参考文献

被引文献

相似文献

元条形码有潜力成为一种快速、灵敏且有效的方法,用于识别复杂环境样本中的物种。物种的准确分子鉴定取决于生成与生物物种相对应的操作分类单元(OTU)的能力。由于有时使用这种方法对生物多样性进行大量估计,因此非常需要测试用于推导 OTU 的数据分析方法的有效性。在这里,我们使用浮游动物物种的模拟群落和自然群落评估了将复杂样本中的长度可变 18S 扩增子聚类到 OTU 中的各种方法的性能。我们比较了由以下组合组成的分析程序:(1) 严格和宽松的数据过滤,(2) 包含和删除的单例序列,(3) 三种常用的聚类算法(mothur、UCLUST 和 UPARSE),以及 (4) 计算序列分歧时处理比对间隙的三种方法。根据所使用方法的组合,模拟群落的 OTU 数量变化了近两个数量级(60-5068 个 OTU),而自然群落的 OTU 数量变化了三个数量级(22-22191 个 OTU)。使用宽松的过滤和包含单例极大地增加了 OTU 数量,但没有增加恢复物种的能力。我们的结果还表明,计算序列分歧时处理缺口的方法会对 OTU 的数量产生很大影响。我们的研究结果与涵盖分类学上不同物种并采用长度变异广泛的 rRNA 基因等标记的研究特别相关。
Metabarcoding has the potential to become a rapid, sensitive, and effective approach for identifying species in complex environmental samples. Accurate molecular identification of species depends on the ability to generate operational taxonomic units (OTUs) that correspond to biological species. Due to the sometimes enormous estimates of biodiversity using this method, there is a great need to test the efficacy of data analysis methods used to derive OTUs. Here, we evaluate the performance of various methods for clustering length variable 18S amplicons from complex samples into OTUs using a mock community and a natural community of zooplankton species. We compare analytic procedures consisting of a combination of (1) stringent and relaxed data filtering, (2) singleton sequences included and removed, (3) three commonly used clustering algorithms (mothur, UCLUST, and UPARSE), and (4) three methods of treating alignment gaps when calculating sequence divergence. Depending on the combination of methods used, the number of OTUs varied by nearly two orders of magnitude for the mock community (60–5068 OTUs) and three orders of magnitude for the natural community (22–22191 OTUs). The use of relaxed filtering and the inclusion of singletons greatly inflated OTU numbers without increasing the ability to recover species. Our results also suggest that the method used to treat gaps when calculating sequence divergence can have a great impact on the number of OTUs. Our findings are particularly relevant to studies that cover taxonomically diverse species and employ markers such as rRNA genes in which length variation is extensive.
DOI: 10.1186/1471-2105-11-38
发表时间: 2010-01-20
期刊: BMC bioinformatics
影响因子: 3
作者:
Falgueras J;Lara AJ;Fernández-Pozo N;Cantón FR;Pérez-Trabado G;Claros MG
通讯作者: Claros MG
DOI: 10.1098/rsbl.2003.0025
发表时间: 2003-08-07
影响因子: 4.7
作者:
Hebert, PDN;Ratnasingham, S;deWaard, JR
通讯作者: deWaard, JR
DOI: 10.1038/ncomms1095
发表时间: 2010-10-19
影响因子: 16.6
作者:
Fonseca, Vera G.;Carvalho, Gary R.;Sung, Way;Johnson, Harriet F.;Power, Deborah M.;Neill, Simon P.;Packer, Margaret;Blaxter, Mark L.;Lambshead, P. John D.;Thomas, W. Kelley;Creer, Simon
通讯作者: Creer, Simon
DOI: 10.1186/gb-2007-8-7-r143
发表时间: 2007
期刊: Genome biology
影响因子: 12.3
作者:
Huse SM;Huber JA;Morrison HG;Sogin ML;Welch DM
通讯作者: Welch DM
DOI: 10.1371/journal.pone.0074371
发表时间: 2013
期刊: PloS one
影响因子: 3.7
作者:
Egge E;Bittner L;Andersen T;Audic S;de Vargas C;Edvardsen B
通讯作者: Edvardsen B