A Bayesian taxonomic classification method for 16S rRNA gene sequences with improved species-level accuracy.

A Bayesian taxonomic classification method for 16S rRNA gene sequences with improved species-level accuracy.
复制标题

DOI:
10.1186/s12859-017-1670-4
复制
发表时间:
2017-05-10
期刊:
影响因子:
3
通讯作者:
Dong Q
Dong Q
中科院分区:
生物学4区
文献类型:
--
作者:
Gao X;Lin H;Revanna K;Dong Q

文献摘要

被引文献

相似文献

16S rRNA基因序列的物种水平分类仍然是微生物组研究人员面临的严峻挑战,因为现有的16S rRNA基因序列分类工具要么不能提供物种水平的分类,要么分类结果不可靠。不可靠的结果是由于现有方法的局限性,这些方法要么缺乏可靠的基于概率的标准来评估其分类分配的置信度,要么使用核苷酸k-mer频率作为序列相似性测量的代理。我们开发了一种方法,该方法比现有方法显着改善了物种水平的分类结果。我们的方法使用成对序列比对计算查询序列和数据库命中之间的真实序列相似性。根据每个查询序列的多个数据库命中的最低共同祖先将分类从种分配到门水平,并通过自举置信度评分进一步评估分类可靠性。该方法的新颖之处在于,每个数据库命中对查询序列的分类分配的贡献由基于数据库命中与查询序列的序列相似度的贝叶斯后验概率加权。我们的方法不需要任何针对不同分类组的特定训练数据集。相反,只需要一个参考数据库来比对查询序列,使我们的方法很容易适用于16S rRNA基因的不同区域或其他系统发育标记基因。对16S rRNA或其他系统发育标记基因进行可靠的物种水平分类对微生物组研究至关重要。我们的软件显示出比现有工具更高的分类准确性,并且我们提供基于概率的置信度评分来评估基于多个数据库匹配查询序列的分类分类分配的可靠性。尽管计算成本较高,但我们的方法仍然适用于实际目的的大规模微生物组数据集分析。此外,该方法可用于任何系统发育标记基因序列的分类分类。我们的软件叫做BLCA,可以在https://github.com/qunfengdong/BLCA上免费获得。
Species-level classification for 16S rRNA gene sequences remains a serious challenge for microbiome researchers, because existing taxonomic classification tools for 16S rRNA gene sequences either do not provide species-level classification, or their classification results are unreliable. The unreliable results are due to the limitations in the existing methods which either lack solid probabilistic-based criteria to evaluate the confidence of their taxonomic assignments, or use nucleotide k-mer frequency as the proxy for sequence similarity measurement. We have developed a method that shows significantly improved species-level classification results over existing methods. Our method calculates true sequence similarity between query sequences and database hits using pairwise sequence alignment. Taxonomic classifications are assigned from the species to the phylum levels based on the lowest common ancestors of multiple database hits for each query sequence, and further classification reliabilities are evaluated by bootstrap confidence scores. The novelty of our method is that the contribution of each database hit to the taxonomic assignment of the query sequence is weighted by a Bayesian posterior probability based upon the degree of sequence similarity of the database hit to the query sequence. Our method does not need any training datasets specific for different taxonomic groups. Instead only a reference database is required for aligning to the query sequences, making our method easily applicable for different regions of the 16S rRNA gene or other phylogenetic marker genes. Reliable species-level classification for 16S rRNA or other phylogenetic marker genes is critical for microbiome research. Our software shows significantly higher classification accuracy than the existing tools and we provide probabilistic-based confidence scores to evaluate the reliability of our taxonomic classification assignments based on multiple database matches to query sequences. Despite its higher computational costs, our method is still suitable for analyzing large-scale microbiome datasets for practical purposes. Furthermore, our method can be applied for taxonomic classification of any phylogenetic marker gene sequences. Our software, called BLCA, is freely available at https://github.com/qunfengdong/BLCA.