MetaVelvet: an extension of Velvet assembler to de novo metagenome assembly from short sequence reads

MetaVelvet: an extension of Velvet assembler to de novo metagenome assembly from short sequence reads
复制标题

DOI:
10.1145/2147805.2147818
复制
发表时间:
2011-08
影响因子:
14.9
通讯作者:
Toshiaki Namiki;Tsuyoshi Hachiya;Hideaki Tanaka;Y. Sakakibara
Toshiaki Namiki;Tsuyoshi Hachiya;Hideaki Tanaka;Y. Sakakibara
中科院分区:
生物学2区
文献类型:
--
作者:
Toshiaki Namiki;Tsuyoshi Hachiya;Hideaki Tanaka;Y. Sakakibara

文献摘要

被引文献

相似文献

动机:“宏基因组学”分析的一个重要步骤是从微生物群落中多个物种的混合序列读取中组装多个基因组。大多数传统管道采用具有仔细优化参数的单基因组组装器,并对所得支架进行后处理以纠正组装错误。使用单基因组组装程序进行从头宏基因组组装的局限性在于,不同物种之间共享的高度保守的序列通常会导致嵌合重叠群,并且高丰度物种的序列可能会被错误地识别为单个基因组中的重复序列,从而导致许多小片段支架。当从非常短的序列读取进行组装时,宏基因组组装问题变得更加困难。方法:我们修改并扩展了一个基于单基因组和 de Bruijn 图的组装程序,称为“Velvet”[27],用于短读取到宏基因组组装,称为“MetaVelvet”,用于多个物种的混合短读取。我们的基本思想是首先将由混合短读段构建的 de Bruijn 图分解为单个子图,然后基于每个分解的 de Bruijn 子图作为分离物种基因组构建支架。我们利用图连通性和覆盖(丰度)差异这两个特征来分解 de Bruijn 图。结果:在模拟数据集上,MetaVelvet 成功地生成了比任何同类单基因组组装程序更高的 N50 分数和更小的嵌合支架,生成高质量的支架以及使用 Velvet 从分离物种序列读取中进行单独组装,并且 MetaVelvet 甚至可以重建覆盖率相对较低的基因组序列作为支架。在人类肠道微生物读取数据的真实数据集上,MetaVelvet 生成了更长的支架,增加了预测基因的数量,并改进了门级分类法的分配,因为无法分配给任何分类法的预测基因的比率降低了。可用性:MetaVelvet 的源代码可在 GNU 通用公共许可证下从 http://metavevet.dna.bio.keio.ac.jp 免费获取。
Motivation: An important step of "metagenomics" analysis is the assembly of multiple genomes from mixed sequence reads of multiple species in a microbial community. Most conventional pipelines employ a single-genome assembler with carefully optimized parameters and post-process the resulting scaffolds to correct assembly errors. Limitations of the use of a single-genome assembler for de novo metagenome assembly are that highly conserved sequences shared between different species often causes chimera contigs, and sequences of highly abundant species are likely mis-identified as repeats in a single genome, resulting in a number of small fragmented scaffolds. The metagenome assembly problem becomes harder when assembling from very short sequence reads. Method: We modified and extended a single-genome and de Bruijn-graph based assembler, known as "Velvet" [27], for short reads to metagenome assembly, called "MetaVelvet", for mixed short reads of multiple species. Our fundamental ideas are first decomposing de Bruijn graph constructed from mixed short reads into individual sub-graphs and second building scaffolds based on every decomposed de Bruijn sub-graph as isolate species genome. We make use of two features, graph connectivity and coverage (abundance) difference, for the decomposition of de Bruijn graph. Results: On simulated datasets, MetaVelvet succeeded to generate higher N50 scores and smaller chimeric scaffolds than any compared single-genome assemblers, produce high-quality scaffolds as well as the separate assembly using Velvet from isolated species sequence reads, and MetaVelvet reconstructed even relatively low-coverage genome sequences as scaffolds. On a real dataset of Human Gut microbial read data, MetaVelvet produced longer scaffolds, increased the number of predicted genes, and improved the assignments of a phylum-level taxonomy in the sense that the rate of predicted genes that cannot be assigned to any tanoxomy is reduced. Availability: The source code of MetaVelvet is freely available at http://metavelvet.dna.bio.keio.ac.jp under the GNU General Public License.