Comparing K-mer based methods for improved classification of 16S sequences.

Comparing K-mer based methods for improved classification of 16S sequences.
复制标题

DOI:
10.1186/s12859-015-0647-4
复制
发表时间:
2015-07-01
期刊:
影响因子:
3
通讯作者:
Snipen L
Snipen L
中科院分区:
生物学4区
文献类型:
--
作者:
Vinje H;Liland KH;Almøy T;Snipen L

文献摘要

参考文献

被引文献

相似文献

对精确和稳定的分类学分类的需求与现代微生物学高度相关。随着可访问序列数据量的爆炸式增长,分类方法的焦点也发生了转变。以前,基于对齐的方法是最适用的工具。现在,就速度和准确性而言,基于滑动窗口计数 K-mers 的方法是最有趣的分类方法。在这里,我们对 16S rRNA 基因的五种不同的基于 K-mer 的分类方法进行了系统比较。这些方法在数据使用和建模策略方面都各不相同。我们的研究基于 RDP 项目中众所周知且常用的朴素贝叶斯分类器,并且在两个不同的数据集、全长序列以及典型读长片段上实施和测试了其他四种方法。这些方法获得的分类误差差异似乎很小,但对于测试的两个数据集来说它们都是稳定的。预处理最近邻 (PLSNN) 方法对于全长 16S rRNA 序列表现最佳,明显优于朴素贝叶斯 RDP 方法。在片段序列上,朴素贝叶斯多项式方法表现最好,明显优于所有其他方法。对于探索的两个数据集,以及全长和片段序列,所有五种方法都达到了误差平台。我们得出的结论是,没有一种基于 K-mer 的方法对于全长序列和片段(reads)的分类来说是普遍最佳的。所有方法都接近误差平台,表明需要改进训练数据来改进这里的分类。对于序列很少的属,分类错误最常见。为了改进分类法和测试新的分类方法,需要更好、更通用和更强大的训练数据集至关重要。
The need for precise and stable taxonomic classification is highly relevant in modern microbiology. Parallel to the explosion in the amount of sequence data accessible, there has also been a shift in focus for classification methods. Previously, alignment-based methods were the most applicable tools. Now, methods based on counting K-mers by sliding windows are the most interesting classification approach with respect to both speed and accuracy. Here, we present a systematic comparison on five different K-mer based classification methods for the 16S rRNA gene. The methods differ from each other both in data usage and modelling strategies. We have based our study on the commonly known and well-used naïve Bayes classifier from the RDP project, and four other methods were implemented and tested on two different data sets, on full-length sequences as well as fragments of typical read-length. The difference in classification error obtained by the methods seemed to be small, but they were stable and for both data sets tested. The Preprocessed nearest-neighbour (PLSNN) method performed best for full-length 16S rRNA sequences, significantly better than the naïve Bayes RDP method. On fragmented sequences the naïve Bayes Multinomial method performed best, significantly better than all other methods. For both data sets explored, and on both full-length and fragmented sequences, all the five methods reached an error-plateau. We conclude that no K-mer based method is universally best for classifying both full-length sequences and fragments (reads). All methods approach an error plateau indicating improved training data is needed to improve classification from here. Classification errors occur most frequent for genera with few sequences present. For improving the taxonomy and testing new classification methods, the need for a better and more universal and robust training data set is crucial.
高度平行的焦索测序者产生的16S rRNA序列的准确分类学分配。
DOI: 10.1093/nar/gkn491
发表时间: 2008-10
影响因子: 14.9
作者:
Liu Z;DeSantis TZ;Andersen GL;Knight R
通讯作者: Knight R
DOI: 10.1186/1471-2105-12-318
发表时间: 2011-08-03
期刊: BMC bioinformatics
影响因子: 3
作者:
Mehmood T;Martens H;Saebø S;Warringer J;Snipen L
通讯作者: Snipen L
DOI: 10.1186/1471-2105-13-97
发表时间: 2012-05-14
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Mehmood, Tahir;Bohlin, Jon;Snipen, Lars
通讯作者: Snipen, Lars
DOI: 10.1093/bioinformatics/18.1.39
发表时间: 2002-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Nguyen, DV;Rocke, DM
通讯作者: Rocke, DM
DOI: 10.1093/nar/gkq873
发表时间: 2010-12
影响因子: 14.9
作者:
Claesson MJ;Wang Q;O'Sullivan O;Greene-Diniz R;Cole JR;Ross RP;O'Toole PW
通讯作者: O'Toole PW