Deep learning models for bacteria taxonomic classification of metagenomic data.

Deep learning models for bacteria taxonomic classification of metagenomic data.
复制标题

细菌分类学分类的深度学习模型。

DOI:
10.1186/s12859-018-2182-6
复制
发表时间:
2018-07-09
期刊:
影响因子:
3
通讯作者:
Urso A
Urso A
中科院分区:
生物学4区
文献类型:
--
作者:
Fiannaca A;La Paglia L;La Rosa M;Lo Bosco G;Renda G;Rizzo R;Gaglio S;Urso A

文献摘要

参考文献

被引文献

相似文献

翻译生物信息学中的一个开放性挑战是分析来自各种环境样品的测序宏基因组。当然,一些研究表明,16 S核糖体RNA可以被认为是细菌属水平分类的条形码,但到目前为止,很难从RNA-seq短读数据中识别宏基因组数据的正确组成。使用两种下一代测序技术,即全基因组鸟枪(WGS)和扩增子(AMP)生成16 S短读段数据;通常,前者经过过滤以获得属于16 S鸟枪(SG)的短读段,而后者仅考虑一些特定的16 S高变区。上述两种测序技术SG和AMP交替使用,因此在这项工作中,我们提出了一种用于宏基因组数据分类的深度学习方法,可以用于这两种技术。为了测试所提出的流水线,我们模拟了来自1000个16 S全长序列的SG和AMP短读段。然后,我们采用k-mer表示将序列作为向量映射到数值空间。最后,我们训练了两种不同的深度学习架构,即,卷积神经网络(CNN)和深度信念网络(DBN),获得每个分类群的训练模型。我们测试了我们提出的方法,以找到最佳参数配置,并将我们的结果与用于细菌识别的参考分类器(称为RDP分类器)提供的分类性能进行了比较。我们优于RDP分类器在每个分类层次与两个架构。例如,在属的水平上,CNN和DBN在AMP短读段的准确率都达到了91.3%,而RDP分类器在相同的数据下获得了83.8%的准确率。在这项工作中,我们提出了一种基于k-mer表示和深度学习架构的16 S短读段序列分类技术,其中每个分类单元(从门到属)都会生成一个分类模型。实验结果证实了所提出的流水线作为一种有效的方法分类细菌序列;因此,我们的方法可以集成到最常见的宏基因组分析工具中。根据所得到的结果,它可以成功地用于分类SG和AMP数据。本文的在线版本(10.1186/s12859-018-2182-6)包含补充材料,可供授权用户使用。
An open challenge in translational bioinformatics is the analysis of sequenced metagenomes from various environmental samples. Of course, several studies demonstrated the 16S ribosomal RNA could be considered as a barcode for bacteria classification at the genus level, but till now it is hard to identify the correct composition of metagenomic data from RNA-seq short-read data. 16S short-read data are generated using two next generation sequencing technologies, i.e. whole genome shotgun (WGS) and amplicon (AMP); typically, the former is filtered to obtain short-reads belonging to a 16S shotgun (SG), whereas the latter take into account only some specific 16S hypervariable regions. The above mentioned two sequencing technologies, SG and AMP, are used alternatively, for this reason in this work we propose a deep learning approach for taxonomic classification of metagenomic data, that can be employed for both of them. To test the proposed pipeline, we simulated both SG and AMP short-reads, from 1000 16S full-length sequences. Then, we adopted a k-mer representation to map sequences as vectors into a numerical space. Finally, we trained two different deep learning architecture, i.e., convolutional neural network (CNN) and deep belief network (DBN), obtaining a trained model for each taxon. We tested our proposed methodology to find the best parameters configuration, and we compared our results against the classification performances provided by a reference classifier for bacteria identification, known as RDP classifier. We outperformed the RDP classifier at each taxonomic level with both architectures. For instance, at the genus level, both CNN and DBN reached 91.3% of accuracy with AMP short-reads, whereas RDP classifier obtained 83.8% with the same data. In this work, we proposed a 16S short-read sequences classification technique based on k-mer representation and deep learning architecture, in which each taxon (from phylum to genus) generates a classification model. Experimental results confirm the proposed pipeline as a valid approach for classifying bacteria sequences; for this reason, our approach could be integrated into the most common tools for metagenomic analysis. According to obtained results, it can be successfully used for classifying both SG and AMP data. The online version of this article (10.1186/s12859-018-2182-6) contains supplementary material, which is available to authorized users.
DOI: 10.1214/aoms/1177729694
发表时间: 1951-01-01
影响因子: --
作者:
KULLBACK, S;LEIBLER, RA
通讯作者: LEIBLER, RA
MOCAT:一种宏基因组组件和基因预测工具包。
DOI: 10.1371/journal.pone.0047656
发表时间: 2012
期刊: PloS one
影响因子: 3.7
作者:
Kultima JR;Sunagawa S;Li J;Chen W;Chen H;Mende DR;Arumugam M;Pan Q;Liu B;Qin J;Wang J;Bork P
通讯作者: Bork P
DOI: 10.1101/gr.5969107
发表时间: 2007-03-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Huson, Daniel H.;Auch, Alexander F.;Schuster, Stephan C.
通讯作者: Schuster, Stephan C.
DOI: 10.1093/bioinformatics/btp157
发表时间: 2009-05-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Nawrocki, Eric P.;Kolbe, Diana L.;Eddy, Sean R.
通讯作者: Eddy, Sean R.
DOI: 10.1093/bib/bbw068
发表时间: 2017-09-01
影响因子: 9.5
作者:
Min, Seonwoo;Lee, Byunghan;Yoon, Sungroh
通讯作者: Yoon, Sungroh