VGEA: an RNA viral assembly toolkit.

VGEA: an RNA viral assembly toolkit.
复制标题

DOI:
10.7717/peerj.12129
复制
发表时间:
2021
期刊:
影响因子:
2.7
通讯作者:
Happi CT
Happi CT
中科院分区:
生物学3区
文献类型:
--
作者:
Oluniyi PE;Ajogbasile F;Oguzie J;Uwanibe J;Kayode A;Happi A;Ugwu A;Olumade T;Ogunsanya O;Eromon PE;Folarin O;Frost SDW;Heeney J;Happi CT

文献摘要

参考文献

被引文献

相似文献

基于下一代测序(NGS)的研究极大地增加了我们对病毒多样性的理解。从NGS实验中获得的病毒序列数据是丰富的信息来源,这些数据可用于研究其流行病学,进化,传播模式,还可为药物和疫苗设计提供信息。然而,由于病毒基因组的高突变率和在同一感染宿主中形成准种,因此对生物信息学提出了巨大的挑战,需要实施先进的生物信息学工具来组装代表个体患者中传播的病毒群体的共有基因组。已经开发了许多工具来预处理测序读数、从头进行或参考辅助病毒基因组组装以及评估所获得的基因组的质量。然而,这些工具中的大多数作为独立的工作流程存在,并且通常需要巨大的计算资源。在这里,我们提出了(病毒基因组容易分析),分析RNA病毒基因组的Snakemake工作流程。VGEA使用户能够将测序读段映射到人类基因组以去除人类污染物,将bam文件拆分为正向和反向读段,进行正向和反向读段的从头组装以生成重叠群,对读段进行质量和污染预处理,使用由用户选择的参考序列补充的校正重叠群将读段映射到针对样品定制的参考,并评估/比较基因组组装。我们设计了一个项目,目的是从现有的/独立的生物信息学工具中创建一个灵活,易于使用和一体化的管道,用于病毒基因组分析,可以在个人计算机上部署。VGEA建立在Snakemake工作流程管理系统之上,并利用现有的工具完成每个步骤:fastp用于读段修剪和读段水平质量控制,BWA用于将测序读段映射到人类参考基因组,SAMtools用于提取未映射的读段并且还用于将bam文件拆分成fastq文件,IVA用于从头组装以生成重叠群,shrift用于预处理读段以获得质量和污染,然后,使用补充有用户选择的现有参考序列的校正的重叠群,用于清洁QUAST的颤抖组装的SeqKit,用于评价/评估基因组组装的质量的QUAST,以及用于聚合来自fastp、BWA和QUAST的结果的MultiQC,映射到针对样品定制的参考。我们的管道成功地测试和验证了SARS-CoV-2(n = 20),HIV-1(n = 20)和拉沙病毒(n = 20)数据集,所有这些数据集都已公开。VGEA可以在GitHub上免费获得:https://github.com/pauloluniyi/VGEA,遵循GNU通用公共许可证。
Next generation sequencing (NGS)-based studies have vastly increased our understanding of viral diversity. Viral sequence data obtained from NGS experiments are a rich source of information, these data can be used to study their epidemiology, evolution, transmission patterns, and can also inform drug and vaccine design. Viral genomes, however, represent a great challenge to bioinformatics due to their high mutation rate and forming quasispecies in the same infected host, bringing about the need to implement advanced bioinformatics tools to assemble consensus genomes well-representative of the viral population circulating in individual patients. Many tools have been developed to preprocess sequencing reads, carry-out de novo or reference-assisted assembly of viral genomes and assess the quality of the genomes obtained. Most of these tools however exist as standalone workflows and usually require huge computational resources. Here we present (Viral Genomes Easily Analyzed), a Snakemake workflow for analyzing RNA viral genomes. VGEA enables users to map sequencing reads to the human genome to remove human contaminants, split bam files into forward and reverse reads, carry out de novo assembly of forward and reverse reads to generate contigs, pre-process reads for quality and contamination, map reads to a reference tailored to the sample using corrected contigs supplemented by the user’s choice of reference sequences and evaluate/compare genome assemblies. We designed a project with the aim of creating a flexible, easy-to-use and all-in-one pipeline from existing/stand-alone bioinformatics tools for viral genome analysis that can be deployed on a personal computer. VGEA was built on the Snakemake workflow management system and utilizes existing tools for each step: fastp for read trimming and read-level quality control, BWA for mapping sequencing reads to the human reference genome, SAMtools for extracting unmapped reads and also for splitting bam files into fastq files, IVA for de novo assembly to generate contigs, shiver to pre-process reads for quality and contamination, then map to a reference tailored to the sample using corrected contigs supplemented with the user’s choice of existing reference sequences, SeqKit for cleaning shiver assembly for QUAST, QUAST to evaluate/assess the quality of genome assemblies and MultiQC for aggregation of the results from fastp, BWA and QUAST. Our pipeline was successfully tested and validated with SARS-CoV-2 (n = 20), HIV-1 (n = 20) and Lassa Virus (n = 20) datasets all of which have been made publicly available. VGEA is freely available on GitHub at: https://github.com/pauloluniyi/VGEA under the GNU General Public License.
DOI: 10.1093/bioinformatics/btab015
发表时间: 2021-07-19
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Posada-Céspedes S;Seifert D;Topolsky I;Jablonski KP;Metzner KJ;Beerenwinkel N
通讯作者: Beerenwinkel N
基因组流行病学揭示了寨卡病毒的多次引入到美国。
DOI: 10.1038/nature22400
发表时间: 2017-06-15
期刊: Nature
影响因子: 64.8
作者:
Grubaugh ND;Ladner JT;Kraemer MUG;Dudas G;Tan AL;Gangavarapu K;Wiley MR;White S;Thézé J;Magnani DM;Prieto K;Reyes D;Bingham AM;Paul LM;Robles-Sikisaka R;Oliveira G;Pronty D;Barcellona CM;Metsky HC;Baniecki ML;Barnes KG;Chak B;Freije CA;Gladden-Young A;Gnirke A;Luo C;MacInnis B;Matranga CB;Park DJ;Qu J;Schaffner SF;Tomkins-Tinch C;West KL;Winnicki SM;Wohl S;Yozwiak NL;Quick J;Fauver JR;Khan K;Brent SE;Reiner RC Jr;Lichtenberger PN;Ricciardi MJ;Bailey VK;Watkins DI;Cone MR;Kopp EW 4th;Hogan KN;Cannons AC;Jean R;Monaghan AJ;Garry RF;Loman NJ;Faria NR;Porcelli MC;Vasquez C;Nagle ER;Cummings DAT;Stanek D;Rambaut A;Sanchez-Lockhart M;Sabeti PC;Gillis LD;Michael SF;Bedford T;Pybus OG;Isern S;Palacios G;Andersen KG
通讯作者: Andersen KG
DOI: 10.1016/s0140-6736(20)30183-5
发表时间: 2020-02-15
期刊: LANCET
影响因子: 168.9
作者:
Huang, Chaolin;Wang, Yeming;Cao, Bin
通讯作者: Cao, Bin
DOI: 10.1086/338820
发表时间: 2002-05-01
影响因子: 11.8
作者:
Chan, PKS
通讯作者: Chan, PKS
DOI: 10.1038/nature22402
发表时间: 2017-06-15
期刊: NATURE
影响因子: 64.8
作者:
Metsky, Hayden C.;Matranga, Christian B.;Sabeti, Pardis C.
通讯作者: Sabeti, Pardis C.