A total crapshoot? Evaluating bioinformatic decisions in animal diet metabarcoding analyses.

A total crapshoot? Evaluating bioinformatic decisions in animal diet metabarcoding analyses.
复制标题

DOI:
10.1002/ece3.6594
复制
发表时间:
2020-09
影响因子:
2.6
通讯作者:
Foster JT
Foster JT
中科院分区:
生物学2区
文献类型:
--
作者:
O'Rourke DR;Bokulich NA;Jusino MA;MacManes MD;Foster JT

文献摘要

参考文献

被引文献

相似文献

元条码研究为自然界混合群落中生物多样性和丰度的估计提供了一种强有力的方法。虽然存在优化样品和序列文库制备的策略,但在动物饮食研究中缺乏扩增子序列数据的生物信息学处理的最佳实践。在这里,我们评估如何在核心生物信息学过程中,包括序列过滤,数据库设计和分类,可以影响动物元编码结果的决定。我们表明,与传统的聚类方法相比,去噪方法具有较低的错误率,尽管通过去除低丰度序列变体,这些差异在很大程度上得到了缓解。我们还发现,从GenBank和BOLD的动物标记基因细胞色素氧化酶I(COI)的可用参考数据集可以是互补的,我们讨论的方法,以改善现有的数据库,包括版本发布。分类学的分类方法可以显著影响结果。例如,与所有其他分类器(vsearch-SINTAX和q2-feature-classifier的BLAST + LCA、Vsearch + LCA和Naive Bayes分类器)相比,常用的生命条形码数据库(BOLD)分类API为使用模拟群落和蝙蝠粪便样本的从物种级别到顺序的样本分配了更少的名称。在生物信息学最佳做法方面缺乏共识,限制了研究之间的比较,并可能带来偏见。我们的工作表明,生物模拟社区提供了一个有用的标准,以评估无数的计算决策影响动物元条码的准确性。此外,这些比较突出了持续评估的必要性,因为新工具的采用,以确保得出的推论反映了有意义的生物学,而不是数字文物。现代动物饮食元巴编码实验是许多替代生物信息学管道的复杂组合。我们选择研究在任何饮食分析中常见的三个核心方面(聚类/去噪,分类和数据库开发)的参数选择如何影响随后的推断。我们的研究结果概述了一套推荐的过程和模板,为未来的评估,新的生物信息学工具变得可用。
Metabarcoding studies provide a powerful approach to estimate the diversity and abundance of organisms in mixed communities in nature. While strategies exist for optimizing sample and sequence library preparation, best practices for bioinformatic processing of amplicon sequence data are lacking in animal diet studies. Here we evaluate how decisions made in core bioinformatic processes, including sequence filtering, database design, and classification, can influence animal metabarcoding results. We show that denoising methods have lower error rates compared to traditional clustering methods, although these differences are largely mitigated by removing low‐abundance sequence variants. We also found that available reference datasets from GenBank and BOLD for the animal marker gene cytochrome oxidase I (COI) can be complementary, and we discuss methods to improve existing databases to include versioned releases. Taxonomic classification methods can dramatically affect results. For example, the commonly used Barcode of Life Database (BOLD) Classification API assigned fewer names to samples from order through species levels using both a mock community and bat guano samples compared to all other classifiers (vsearch‐SINTAX and q2‐feature‐classifier's BLAST + LCA, VSEARCH + LCA, and Naive Bayes classifiers). The lack of consensus on bioinformatics best practices limits comparisons among studies and may introduce biases. Our work suggests that biological mock communities offer a useful standard to evaluate the myriad computational decisions impacting animal metabarcoding accuracy. Further, these comparisons highlight the need for continual evaluations as new tools are adopted to ensure that the inferences drawn reflect meaningful biology instead of digital artifacts. Modern animal diet metabarcoding experiments are complicated assortments of many alternative bioinformatic pipelines. We chose to investigate how parameter choices in three core aspects common to any diet analysis (clustering/denoising, classification, and database development) may influence subsequent inferences. Our findings outline both a set of recommended processes and a template for future evaluations as new bioinformatic tools become available.
DOI: 10.1128/msystems.00062-16
发表时间: 2016-09-01
期刊: MSYSTEMS
影响因子: 6.4
作者:
Bokulich, Nicholas A.;Rideout, Jai Ram;Caporaso, J. Gregory
通讯作者: Caporaso, J. Gregory
DOI: 10.1038/nmeth.3869
发表时间: 2016-07
期刊: Nature methods
影响因子: 48
作者:
Callahan BJ;McMurdie PJ;Rosen MJ;Han AW;Johnson AJ;Holmes SP
通讯作者: Holmes SP
DOI: 10.1111/mec.14734
发表时间: 2019-01
期刊: Molecular ecology
影响因子: 4.9
作者:
Deagle BE;Thomas AC;McInnes JC;Clarke LJ;Vesterinen EJ;Clare EL;Kartzinel TR;Eveson JP
通讯作者: Eveson JP
DOI: 10.1093/gigascience/giy054
发表时间: 2018-05-01
期刊: GigaScience
影响因子: 9.2
作者:
Almeida A;Mitchell AL;Tarkowska A;Finn RD
通讯作者: Finn RD
DOI: 10.1093/nar/gkl986
发表时间: 2010-01-01
影响因子: 14.9
作者:
Benson, Dennis A.;Karsch-Mizrachi, Ilene;Sayers, Eric W.
通讯作者: Sayers, Eric W.