Evaluation of shotgun metagenomics sequence classification methods using in silico and in vitro simulated communities.

Evaluation of shotgun metagenomics sequence classification methods using in silico and in vitro simulated communities.
复制标题

DOI:
10.1186/s12859-015-0788-5
复制
发表时间:
2015-11-04
期刊:
影响因子:
3
通讯作者:
Brinkman FS
Brinkman FS
中科院分区:
生物学4区
文献类型:
--
作者:
Peabody MA;Van Rossum T;Lo R;Brinkman FS

文献摘要

参考文献

被引文献

相似文献

随着许多生物信息学分析方法的发展,元基因组学(研究直接从环境中回收的遗传物质)领域迅速发展。为了确保适当使用这些方法,需要对其准确性和特点进行强有力的比较评估。对于序列读数的分类,这种评估应该包括使用分支排除,当任何参考数据库中没有相同的序列时,这种评估方法更好地评估方法的准确性,这在元基因组分析中是常见的。到目前为止,已经进行了相对较小的评估,诸如支派排除等评估方法仅限于特定方法的作者对新方法的评估。需要的是在多种主要方法之间进行严格、独立的比较,在硅胶和体外测试数据集中使用相同的方法,无论是否使用分支排除等方法,以更好地表征不同条件下的准确性。概述了38种生物信息学方法的特点,重点评估了11个程序的准确性,这些程序具有可以修改的参考数据库,因此在分支排除情况下进行了最稳健的评估。使用硅胶和体外模拟细菌群落对序列读数的分类进行评估。在从物种到纲的分类水平上使用了分支排除--确定方法在逐渐变得更困难的情况下的表现如何。在敏感性、精密度、总体精确度和被评估程序的计算需求方面发现了广泛的变异性。在蒸馏水中只添加了11种细菌的实验中,最受欢迎的程序经常会错误地预测数十到数百种细菌。每种方法的不同特点(是否强制预测等)进行了总结,并讨论了其他分析注意事项。散弹枪元基因组学分类方法的准确性差异很大。在所有评估方案中,没有一种方案明显优于其他方案;相反,结果说明了不同方法针对不同目的的优势。研究人员必须认识到方法的差异,选择最适合他们特定分析的程序,以避免非常误导的结果。鼓励使用标准化数据集进行方法比较,也鼓励使用适用于特定元基因组分析的模拟微生物群落对照。本文的在线版本(doi:10.1186/s12859-0150788-5)包含补充材料,授权用户可以使用。
The field of metagenomics (study of genetic material recovered directly from an environment) has grown rapidly, with many bioinformatics analysis methods being developed. To ensure appropriate use of such methods, robust comparative evaluation of their accuracy and features is needed. For taxonomic classification of sequence reads, such evaluation should include use of clade exclusion, which better evaluates a method’s accuracy when identical sequences are not present in any reference database, as is common in metagenomic analysis. To date, relatively small evaluations have been performed, with evaluation approaches like clade exclusion limited to assessment of new methods by the authors of the given method. What is needed is a rigorous, independent comparison between multiple major methods, using the same in silico and in vitro test datasets, with and without approaches like clade exclusion, to better characterize accuracy under different conditions. An overview of the features of 38 bioinformatics methods is provided, evaluating accuracy with a focus on 11 programs that have reference databases that can be modified and therefore most robustly evaluated with clade exclusion. Taxonomic classification of sequence reads was evaluated using both in silico and in vitro mock bacterial communities. Clade exclusion was used at taxonomic levels from species to class—identifying how well methods perform in progressively more difficult scenarios. A wide range of variability was found in the sensitivity, precision, overall accuracy, and computational demand for the programs evaluated. In experiments where distilled water was spiked with only 11 bacterial species, frequently dozens to hundreds of species were falsely predicted by the most popular programs. The different features of each method (forces predictions or not, etc.) are summarized, and additional analysis considerations discussed. The accuracy of shotgun metagenomics classification methods varies widely. No one program clearly outperformed others in all evaluation scenarios; rather, the results illustrate the strengths of different methods for different purposes. Researchers must appreciate method differences, choosing the program best suited for their particular analysis to avoid very misleading results. Use of standardized datasets for method comparisons is encouraged, as is use of mock microbial community controls suitable for a particular metagenomic analysis. The online version of this article (doi:10.1186/s12859-015-0788-5) contains supplementary material, which is available to authorized users.
DOI: 10.6026/97320630004046
发表时间: 2009-08-20
期刊: Bioinformation
影响因子: 1.9
作者:
Yu F;Sun Y;Liu L;Farmerie W
通讯作者: Farmerie W