Comprehensive benchmarking and ensemble approaches for metagenomic classifiers.
Comprehensive benchmarking and ensemble approaches for metagenomic classifiers.
复制标题
DOI:
10.1186/s13059-017-1299-7
复制
发表时间:
2017-09-21
期刊:
影响因子:
12.3
通讯作者:
Mason CE
中科院分区:
文献类型:
--
作者:
McIntyre ABR;Ounit R;Afshinnekoo E;Prill RJ;Hénaff E;Alexander N;Minot SS;Danko D;Foox J;Ahsanuddin S;Tighe S;Hasan NA;Subramanian P;Moffat K;Levy S;Lonardi S;Greenfield N;Colwell RR;Rosen GL;Mason CE
One of the main challenges in metagenomics is the identification of microorganisms in clinical and environmental samples. While an extensive and heterogeneous set of computational tools is available to classify microorganisms using whole-genome shotgun sequencing data, comprehensive comparisons of these methods are limited. In this study, we use the largest-to-date set of laboratory-generated and simulated controls across 846 species to evaluate the performance of 11 metagenomic classifiers. Tools were characterized on the basis of their ability to identify taxa at the genus, species, and strain levels, quantify relative abundances of taxa, and classify individual reads to the species level. Strikingly, the number of species identified by the 11 tools can differ by over three orders of magnitude on the same datasets. Various strategies can ameliorate taxonomic misclassification, including abundance filtering, ensemble approaches, and tool intersection. Nevertheless, these strategies were often insufficient to completely eliminate false positives from environmental samples, which are especially important where they concern medically relevant species. Overall, pairing tools with different classification strategies (k-mer, alignment, marker) can combine their respective advantages. This study provides positive and negative controls, titrated standards, and a guide for selecting tools for metagenomic analyses by comparing ranges of precision, accuracy, and recall. We show that proper experimental design and analysis parameters can reduce false positives, provide greater resolution of species in complex metagenomic samples, and improve the interpretation of results. The online version of this article (doi:10.1186/s13059-017-1299-7) contains supplementary material, which is available to authorized users.
登录
查看更多内容
影响因子:
7
作者:
Erlich Y
通讯作者:
Erlich Y
影响因子:
16.6
作者:
Bradley P;Gordon NC;Walker TM;Dunn L;Heys S;Huang B;Earle S;Pankhurst LJ;Anson L;de Cesare M;Piazza P;Votintseva AA;Golubchik T;Wilson DJ;Wyllie DH;Diel R;Niemann S;Feuerriegel S;Kohl TA;Ismail N;Omar SV;Smith EG;Buck D;McVean G;Walker AS;Peto TE;Crook DW;Iqbal Z
通讯作者:
Iqbal Z
影响因子:
12.3
作者:
Koren S;Harhay GP;Smith TP;Bono JL;Harhay DM;Mcvey SD;Radune D;Bergman NH;Phillippy AM
通讯作者:
Phillippy AM
DOI:
10.1093/bioinformatics/btt389
发表时间:
2013-09-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
Ames SK;Hysom DA;Gardner SN;Lloyd GS;Gokhale MB;Allen JE
通讯作者:
Allen JE
影响因子:
4.6
作者:
Karlsson E;Lärkeryd A;Sjödin A;Forsman M;Stenberg P
通讯作者:
Stenberg P