DOE JGI Metagenome Workflow.

DOE JGI Metagenome Workflow.
复制标题

DOI:
10.1128/msystems.00804-20
复制
发表时间:
2021-05-18
期刊:
影响因子:
6.4
通讯作者:
Ivanova NN
Ivanova NN
中科院分区:
生物学2区
文献类型:
--
作者:
Clum A;Huntemann M;Bushnell B;Foster B;Foster B;Roux S;Hajek PP;Varghese N;Mukherjee S;Reddy TBK;Daum C;Yoshinaga Y;O'Malley R;Seshadri R;Kyrpides NC;Eloe-Fadrosh EA;Chen IA;Copeland A;Ivanova NN

文献摘要

被引文献

相似文献

美国能源部联合基因组研究所(JGI)宏基因组工作流执行宏基因组数据处理,包括组装;结构、功能和分类注释;元基因组数据集的合并,这些数据集随后被纳入整合微生物基因组和微生物组(IMG/M) (i -M)。A. Chen, K. Chu, K. Palaniappan, A. Ratner等,核酸研究,49:D751-D763, 2021, https://doi.org/10.1093/nar/gkaa939)比较分析系统,并提供通过JGI数据门户(https://genome.jgi.doe.gov/portal/)下载。该工作流程可扩展到每年运行数千个宏基因组样本,这可能因微生物群落的复杂性和测序深度而异。在这里,我们描述了在工作流程的不同步骤中使用的不同工具、数据库和参数,以帮助解释IMG中可用的宏基因组数据,并使研究人员能够将此工作流程应用于他们自己的数据。我们使用20个公开的沉积物宏基因组来说明不同步骤的计算要求,并突出了数据处理的典型结果。用于读取过滤和宏基因组组装的工作流模块可作为工作流描述语言(WDL)文件(https://code.jgi.doe.gov/BFoster/jgi_meta_wdl)获得。注释和分组工作流模块作为服务提供给用户社区,网址为https://img.jgi.doe.gov/submit,需要在基因组在线数据库(GOLD)中填写项目和相关元数据描述(S. Mukherjee, D. Stamatis, J. Bertsch, G. Ovchinnikova等,Nucleic Acids Res, 49: D723-D733, 2021, https://doi.org/10.1093/nar/gkaa983)。DOE JGI宏基因组工作流是为处理从Illumina fastq文件开始的宏基因组数据集而设计的。它执行数据预处理、纠错、汇编、结构和功能注释以及分组。处理结果以几种标准格式提供,如fasta和gff,并可用于随后整合到集成微生物基因组和微生物组(IMG/M)系统中,在该系统中,它们可以与一组全面的公开可用的宏基因组进行比较。截至2020年7月30日,美国能源部JGI宏基因组工作流程已经处理了7155个JGI宏基因组。在这里,我们提出了一个由JGI开发的宏基因组工作流程,该流程以标准格式生成丰富的数据,并已优化用于下游分析,从微生物群落的功能和分类组成评估到基因组解析宏基因组学以及新分类群的鉴定和表征。该工作流程目前被用于以一致和标准化的方式分析数千个宏基因组数据集。
The DOE Joint Genome Institute (JGI) Metagenome Workflow performs metagenome data processing, including assembly; structural, functional, and taxonomic annotation; and binning of metagenomic data sets that are subsequently included into the Integrated Microbial Genomes and Microbiomes (IMG/M) (I.-M. A. Chen, K. Chu, K. Palaniappan, A. Ratner, et al., Nucleic Acids Res, 49:D751–D763, 2021, https://doi.org/10.1093/nar/gkaa939) comparative analysis system and provided for download via the JGI data portal (https://genome.jgi.doe.gov/portal/). This workflow scales to run on thousands of metagenome samples per year, which can vary by the complexity of microbial communities and sequencing depth. Here, we describe the different tools, databases, and parameters used at different steps of the workflow to help with the interpretation of metagenome data available in IMG and to enable researchers to apply this workflow to their own data. We use 20 publicly available sediment metagenomes to illustrate the computing requirements for the different steps and highlight the typical results of data processing. The workflow modules for read filtering and metagenome assembly are available as a workflow description language (WDL) file (https://code.jgi.doe.gov/BFoster/jgi_meta_wdl). The workflow modules for annotation and binning are provided as a service to the user community at https://img.jgi.doe.gov/submit and require filling out the project and associated metadata descriptions in the Genomes OnLine Database (GOLD) (S. Mukherjee, D. Stamatis, J. Bertsch, G. Ovchinnikova, et al., Nucleic Acids Res, 49:D723–D733, 2021, https://doi.org/10.1093/nar/gkaa983). IMPORTANCE The DOE JGI Metagenome Workflow is designed for processing metagenomic data sets starting from Illumina fastq files. It performs data preprocessing, error correction, assembly, structural and functional annotation, and binning. The results of processing are provided in several standard formats, such as fasta and gff, and can be used for subsequent integration into the Integrated Microbial Genomes and Microbiomes (IMG/M) system where they can be compared to a comprehensive set of publicly available metagenomes. As of 30 July 2020, 7,155 JGI metagenomes have been processed by the DOE JGI Metagenome Workflow. Here, we present a metagenome workflow developed at the JGI that generates rich data in standard formats and has been optimized for downstream analyses ranging from assessment of the functional and taxonomic composition of microbial communities to genome-resolved metagenomics and the identification and characterization of novel taxa. This workflow is currently being used to analyze thousands of metagenomic data sets in a consistent and standardized manner.