A comparative evaluation of sequence classification programs

A comparative evaluation of sequence classification programs
复制标题

DOI:
10.1186/1471-2105-13-92
复制
发表时间:
2012-05-10
期刊:
影响因子:
3
通讯作者:
Cummings, Michael P.
Cummings, Michael P.
中科院分区:
生物学4区
文献类型:
--
作者:
Bazinet, Adam L.;Cummings, Michael P.

文献摘要

被引文献

相似文献

背景资料:现代基因组学中的一个基本问题是对来自环境采样的DNA序列片段进行分类或功能分类(即,宏基因组学)。已经提出了几种不同的方法来有效地和高效地做到这一点,并且许多方法已经在软件中实现。除了改变其基本的分类算法外,一些方法还筛选“条形码基因”(如16S rRNA或各种类型的蛋白质编码基因)的序列读数。由于方法的数量和复杂性,它可以是很难为研究人员选择一个是非常适合于特定的analysis.Results:我们分为非常大的数量的程序,近年来已发布的解决序列分类问题的一般算法的基础上,他们使用的比较查询序列对数据库的序列分为三个主要类别。我们还评估了其分类和功能组成的数据集上的每个类别中的领先程序的性能是know.Conclusions:我们发现显着的变异,在分类准确性,精度和资源消耗的序列分类程序时,用于分析各种宏基因组学数据集。然而,我们观察到一些一般的趋势和模式,将是有用的研究人员使用序列分类程序。
Background: A fundamental problem in modern genomics is to taxonomically or functionally classify DNA sequence fragments derived from environmental sampling (i.e., metagenomics). Several different methods have been proposed for doing this effectively and efficiently, and many have been implemented in software. In addition to varying their basic algorithmic approach to classification, some methods screen sequence reads for 'barcoding genes' like 16S rRNA, or various types of protein-coding genes. Due to the sheer number and complexity of methods, it can be difficult for a researcher to choose one that is well-suited for a particular analysis.Results: We divided the very large number of programs that have been released in recent years for solving the sequence classification problem into three main categories based on the general algorithm they use to compare a query sequence against a database of sequences. We also evaluated the performance of the leading programs in each category on data sets whose taxonomic and functional composition is known.Conclusions: We found significant variability in classification accuracy, precision, and resource consumption of sequence classification programs when used to analyze various metagenomics data sets. However, we observe some general trends and patterns that will be useful to researchers who use sequence classification programs.