A total crapshoot? Evaluating bioinformatic decisions in animal diet metabarcoding analyses.
A total crapshoot? Evaluating bioinformatic decisions in animal diet metabarcoding analyses.
复制标题
DOI:
10.1002/ece3.6594
复制
发表时间:
2020-09
影响因子:
2.6
通讯作者:
Foster JT
中科院分区:
文献类型:
--
作者:
O'Rourke DR;Bokulich NA;Jusino MA;MacManes MD;Foster JT
Metabarcoding studies provide a powerful approach to estimate the diversity and abundance of organisms in mixed communities in nature. While strategies exist for optimizing sample and sequence library preparation, best practices for bioinformatic processing of amplicon sequence data are lacking in animal diet studies. Here we evaluate how decisions made in core bioinformatic processes, including sequence filtering, database design, and classification, can influence animal metabarcoding results. We show that denoising methods have lower error rates compared to traditional clustering methods, although these differences are largely mitigated by removing low‐abundance sequence variants. We also found that available reference datasets from GenBank and BOLD for the animal marker gene cytochrome oxidase I (COI) can be complementary, and we discuss methods to improve existing databases to include versioned releases. Taxonomic classification methods can dramatically affect results. For example, the commonly used Barcode of Life Database (BOLD) Classification API assigned fewer names to samples from order through species levels using both a mock community and bat guano samples compared to all other classifiers (vsearch‐SINTAX and q2‐feature‐classifier's BLAST + LCA, VSEARCH + LCA, and Naive Bayes classifiers). The lack of consensus on bioinformatics best practices limits comparisons among studies and may introduce biases. Our work suggests that biological mock communities offer a useful standard to evaluate the myriad computational decisions impacting animal metabarcoding accuracy. Further, these comparisons highlight the need for continual evaluations as new tools are adopted to ensure that the inferences drawn reflect meaningful biology instead of digital artifacts. Modern animal diet metabarcoding experiments are complicated assortments of many alternative bioinformatic pipelines. We chose to investigate how parameter choices in three core aspects common to any diet analysis (clustering/denoising, classification, and database development) may influence subsequent inferences. Our findings outline both a set of recommended processes and a template for future evaluations as new bioinformatic tools become available.
登录
查看更多内容
影响因子:
6.4
作者:
Bokulich, Nicholas A.;Rideout, Jai Ram;Caporaso, J. Gregory
通讯作者:
Caporaso, J. Gregory
影响因子:
48
作者:
Callahan BJ;McMurdie PJ;Rosen MJ;Han AW;Johnson AJ;Holmes SP
通讯作者:
Holmes SP
影响因子:
4.9
作者:
Deagle BE;Thomas AC;McInnes JC;Clarke LJ;Vesterinen EJ;Clare EL;Kartzinel TR;Eveson JP
通讯作者:
Eveson JP
影响因子:
9.2
作者:
Almeida A;Mitchell AL;Tarkowska A;Finn RD
通讯作者:
Finn RD
影响因子:
14.9
作者:
Benson, Dennis A.;Karsch-Mizrachi, Ilene;Sayers, Eric W.
通讯作者:
Sayers, Eric W.