MEGAN analysis of metagenomic data

MEGAN analysis of metagenomic data
复制标题

DOI:
10.1101/gr.5969107
复制
发表时间:
2007-03-01
期刊:
影响因子:
7
通讯作者:
Schuster, Stephan C.
Schuster, Stephan C.
中科院分区:
生物学1区
文献类型:
--
作者:
Huson, Daniel H.;Auch, Alexander F.;Schuster, Stephan C.

文献摘要

被引文献

相似文献

宏基因组学是使用目标或随机测序从共同栖息地获得的生物体样本的基因组内容的研究。目标包括了解微生物多样性的程度和作用。这种样本的分类内容通常是通过与已知序列的序列数据库进行比较来估计的。大多数已发表的研究使用了对末端reads、环境fosmid和BAC克隆的完整序列或环境组装的分析。新兴的高通量合成测序技术为低成本的随机“散弹枪”方法铺平了道路。本文介绍了MEGAN,一个新的计算机程序,允许对大型宏基因组数据集进行笔记本电脑分析。在预处理步骤中,使用BLAST或其他比较工具将一组DNA序列与已知序列的数据库进行比较。然后使用MEGAN计算和探索数据集的分类内容,使用NCBI分类法对结果进行总结和排序。一个简单的最低共同祖先算法将读取分配给分类群,这样分配的分类群的分类水平反映了序列的保守程度。该软件允许大型数据集被解剖,而不需要组装或特定的系统发育标记的目标。它提供图形和统计输出,用于比较不同的数据集。该方法应用于几个数据集,包括马尾藻海数据集,最近发表的从猛犸象骨骼中取样的宏基因组数据集,以及几个完整的微生物基因组。此外,还进行了模拟,以评估该方法在不同读取长度下的性能。
Metagenomics is the study of the genomic content of a sample of organisms obtained from a common habitat using targeted or random sequencing. Goals include understanding the extent and role of microbial diversity. The taxonomical content of such a sample is usually estimated by comparison against sequence databases of known sequences. Most published studies use the analysis of paired-end reads, complete sequences of environmental fosmid and BAC clones, or environmental assemblies. Emerging sequencing- by-synthesis technologies with very high throughput are paving the way to low-cost random "shotgun" approaches. This paper introduces MEGAN, a new computer program that allows laptop analysis of large metagenomic data sets. In a preprocessing step, the set of DNA sequences is compared against databases of known sequences using BLAST or another comparison tool. MEGAN is then used to compute and explore the taxonomical content of the data set, employing the NCBI taxonomy to summarize and order the results. A simple lowest common ancestor algorithm assigns reads to taxa such that the taxonomical level of the assigned taxon reflects the level of conservation of the sequence. The software allows large data sets to be dissected without the need for assembly or the targeting of specific phylogenetic markers. It provides graphical and statistical output for comparing different data sets. The approach is applied to several data sets, including the Sargasso Sea data set, a recently published metagenomic data set sampled from a mammoth bone, and several complete microbial genomes. Also, simulations that evaluate the performance of the approach for different read lengths are presented.