Tackling soil diversity with the assembly of large, complex metagenomes

Tackling soil diversity with the assembly of large, complex metagenomes
复制标题

DOI:
10.1073/pnas.1402564111
复制
发表时间:
2014-04-01
影响因子:
11.1
通讯作者:
Brown, C. Titus
Brown, C. Titus
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Howe, Adina Chuang;Jansson, Janet K.;Brown, C. Titus

文献摘要

被引文献

相似文献

深入采样复杂环境中的微生物群落所需的大量测序数据对序列分析提出了新的挑战。从头宏基因组组装有效地减少了待分析的数据总量,但需要大量的计算资源。我们结合联合收割机两个预组装过滤方法-数字归一化和分区-生成以前棘手的大型宏基因组组件。使用人类肠道模拟社区数据集,我们证明了这些方法的结果几乎相同的组件从未处理的数据。然后,我们从匹配的爱荷华州玉米和原生草原土壤中组装了两个总计3980亿bp(相当于88,000个大肠杆菌基因组)的大型土壤宏基因组。利用京都基因和基因组同源性数据库,所得到的组装重叠群可用于鉴定已知代谢途径的分子相互作用和反应网络。尽管如此,超过60%的预测蛋白质的组装不能对已知的数据库进行注释。许多这些未知的蛋白质丰富的玉米和草原土壤中,突出的好处,发现和表征新奇的土壤生物多样性的组装。此外,80%的测序数据无法组装,因为覆盖率低,这表明需要更多的测序数据来表征土壤的功能含量。
The large volumes of sequencing data required to sample deeply the microbial communities of complex environments pose new challenges to sequence analysis. De novo metagenomic assembly effectively reduces the total amount of data to be analyzed but requires substantial computational resources. We combine two preassembly filtering approaches-digital normalization and partitioning-to generate previously intractable large metagenome assemblies. Using a human-gut mock communitydataset, we demonstrate that these methods result in assemblies nearly identical to assemblies from unprocessed data. We then assemble two large soil metagenomes totaling 398 billion bp (equivalent to 88,000 Escherichia coli genomes) from matched Iowa corn and native prairie soils. The resulting assembled contigs could be used to identify molecular interactions and reaction networks of known metabolic pathways using the Kyoto Encyclopedia of Genes and Genomes Orthology database. Nonetheless, more than 60% of predicted proteins in assemblies could not be annotated against known databases. Many of these unknown proteins were abundant in both corn and prairie soils, highlighting the benefits of assembly for the discovery and characterization of novelty in soil biodiversity. Moreover, 80% of the sequencing data could not be assembled because of low coverage, suggesting that considerably more sequencing data are needed to characterize the functional content of soil.