Amino acid based de Bruijn graph algorithm for identifying complete coding genes from metagenomic and metatranscriptomic short reads

Amino acid based de Bruijn graph algorithm for identifying complete coding genes from metagenomic and metatranscriptomic short reads
复制标题

基于氨基酸的 de Bruijn 图算法,用于从宏基因组和宏转录组短读中识别完整编码基因

DOI:
10.1093/nar/gkz017
复制
发表时间:
2019-01
影响因子:
14.9
通讯作者:
Qi Ji
Qi Ji
中科院分区:
生物学2区
文献类型:
--
作者:
Liu Jiemeng;Lian Qichao;Chen Yamao;Qi Ji

文献摘要

参考文献

被引文献

相似文献

摘要下一代测序(NGS)技术的快速发展极大地促进了元基因组研究,揭示了微生物群落的复杂结构及其与环境的相互作用。由于大多数微生物缺乏基因组序列的信息,从头开始组装原核生物基因组,目的是从各种代谢途径中检索完整的编码基因。微生物组成的复杂性和处理海量元基因组数据的负担,给有效和高效的生物信息工具的开发带来了巨大的挑战。在这里,我们提出了一个蛋白质组装器(MetaPA),它基于De Bruijn图搜索寡肽空间,可以应用于元基因组和元翻译测序数据。当涉及公共同源蛋白质序列来指导组装过程时,MetaPA在真实的高通量测序数据集上以83%的高精度将85%的蛋白质组装成完整的序列。MetaPA在元转录数据中的应用成功地识别了相关研究中验证的大多数活跃转录基因。这些结果表明,MetaPA在元基因组和后转录组学研究中都有很好的潜力来表征微生物区系的组成和丰度。
Abstract Metagenomic studies, greatly promoted by the fast development of next-generation sequencing (NGS) technologies, uncover complex structures of microbial communities and their interactions with environment. As the majority of microbes lack information of genome sequences, it is essential to assemble prokaryotic genomes ab initio aiming to retrieve complete coding genes from various metabolic pathways. The complex nature of microbial composition and the burden of handling a vast amount of metagenomic data, bring great challenges to the development of effective and efficient bioinformatic tools. Here we present a protein assembler (MetaPA), based on de Bruijn graph searching on oligopeptide spaces and can be applied on both metagenomic and metatranscriptomic sequencing data. When public homologous protein sequences are involved to guide the assembling procedures, MetaPA assembles 85% of total proteins in complete sequences with high precision of 83% on real high-throughput sequencing datasets. Application of MetaPA on metatranscriptomic data successfully identifies the majority of actively transcribed genes validated in related studies. The results suggest that MetaPA has a good potential in both metagenomic and metatranscriptomic studies to characterize the composition and abundance of microbiota.
DOI: 10.1111/1462-2920.12086
发表时间: 2013-06
影响因子: 5.1
作者:
Shakya M;Quince C;Campbell JH;Yang ZK;Schadt CW;Podar M
通讯作者: Podar M
DOI: 10.1126/science.1093857
发表时间: 2004-04-02
期刊: SCIENCE
影响因子: 56.9
作者:
Venter, JC;Remington, K;Smith, HO
通讯作者: Smith, HO
DOI: 10.1038/nature08821
发表时间: 2010-03-04
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
MetaSort 通过降低微生物群落复杂性来理清宏基因组组装
DOI: 10.1038/ncomms14306
发表时间: 2017-01-23
影响因子: 16.6
作者:
Ji P;Zhang Y;Wang J;Zhao F
通讯作者: Zhao F
DOI: 10.1186/gb-2012-13-3-r23
发表时间: 2012
期刊: Genome biology
影响因子: 12.3
作者:
Giannoukos G;Ciulla DM;Huang K;Haas BJ;Izard J;Levin JZ;Livny J;Earl AM;Gevers D;Ward DV;Nusbaum C;Birren BW;Gnirke A
通讯作者: Gnirke A