Use of simulated data sets to evaluate the fidelity of metagenomic processing methods

Use of simulated data sets to evaluate the fidelity of metagenomic processing methods
复制标题

DOI:
10.1038/nmeth1043
复制
发表时间:
2007-06-01
期刊:
影响因子:
48
通讯作者:
Kyrpides, Nikos C.
Kyrpides, Nikos C.
中科院分区:
生物学1区
文献类型:
--
作者:
Mavromatis, Konstantinos;Ivanova, Natalia;Kyrpides, Nikos C.

文献摘要

被引文献

相似文献

宏基因组学是用于研究微生物群落的快速新兴研究领域。为了评估目前用于处理宏基因组序列的方法,我们通过结合从113个分离株基因组随机选择的测序读取来构建了三个模拟的数据集。这些数据集的设计旨在根据复杂性和系统发育组成对真实的宏基因组进行建模。我们使用三个常用的基因组组装器(Phrap,Arachne和Jazz)组装了采样读物,并使用两个流行的基因捕获管道(Fgenesb和Critista/Critista/Crivera/Climmer)预测基因。使用一种基于序列相似性(BLAST HIT分布)和两个基于序列组成(系统植物,寡核苷酸频率)分配方法预测组装重叠群的系统发育起源。我们探索了模拟的社区结构和方法组合对每个处理步骤的保真度的影响,与相应的分离株基因组相比。模拟数据集可在线获得,以促进对元基因组分析的工具的标准化基准测试。
Metagenomics is a rapidly emerging field of research for studying microbial communities. To evaluate methods presently used to process metagenomic sequences, we constructed three simulated data sets of varying complexity by combining sequencing reads randomly selected from 113 isolate genomes. These data sets were designed to model real metagenomes in terms of complexity and phylogenetic composition. We assembled sampled reads using three commonly used genome assemblers (Phrap, Arachne and JAZZ), and predicted genes using two popular gene-finding pipelines (fgenesb and CRITICA/GLIMMER). The phylogenetic origins of the assembled contigs were predicted using one sequence similarity-based ( blast hit distribution) and two sequence composition-based (PhyloPythia, oligonucleotide frequencies) binning methods. We explored the effects of the simulated community structure and method combinations on the fidelity of each processing step by comparison to the corresponding isolate genomes. The simulated data sets are available online to facilitate standardized benchmarking of tools for metagenomic analysis.