xGAP: a python based efficient, modular, extensible and fault tolerant genomic analysis pipeline for variant discovery.

xGAP: a python based efficient, modular, extensible and fault tolerant genomic analysis pipeline for variant discovery.
复制标题

xGAP:一个基于 python 的高效、模块化、可扩展和容错的基因组分析管道,用于变异发现。

DOI:
10.1093/bioinformatics/btaa1097
复制
发表时间:
2021
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Sul,JaeHoon
Sul,JaeHoon
中科院分区:
--
文献类型:
--
作者:
Gorla,Aditya;Jew,Brandon;Zhang,Luke;Sul,JaeHoon

文献摘要

相似文献

自2001年第一个人类基因组测序以来,为研究和临床研究处理和分析下一代测序(NGS)数据的生物信息学方法数量迅速增长,旨在识别影响疾病和特征的遗传变异。为了实现这一目标,首先需要从NGS数据中调用基因变体,这需要多个计算密集型分析步骤。遗憾的是,目前还缺乏一种能够以一种完全自动化、高效、快速、可扩展、模块化、用户友好和容错的方式对NGS数据执行所有这些步骤的开源管道。为解决这一问题,我们引入了一种可扩展的基因组分析流水线xGAP,它利用上述功能实现了改进的GATK最佳实践来分析DNA-SEQ数据。结果xGAP通过将基因组分割成多个更小的区域来实现大规模并行化,从而实现了高可伸缩性。它可以在∼90分钟内处理30倍覆盖的全基因组测序数据。就已发现的变异的准确性而言,xGAP在七个基准WGS数据集上的单核苷酸变异的F1平均得分为99.37%,插入/缺失的平均F1得分为99.20%。我们在多个内部(SGE和SLURM)高性能群集上实现了高度一致的结果。与Churchill管道相比,在类似并行化的情况下,xGAP在亚马逊Web服务上分析50倍覆盖的WGS时速度快20%。最后,xGAP是用户友好和容错的,它可以自动重新启动失败的过程,以最大限度地减少所需的用户干预。可用性和实施xGAP可在https://github.com/Adigorla/xgap.Supplementary信息上获得补充数据可在生物信息学在线上获得。
MotivationSince the first human genome was sequenced in 2001, there has been a rapid growth in the number of bioinformatic methods to process and analyze next-generation sequencing (NGS) data for research and clinical studies that aim to identify genetic variants influencing diseases and traits. To achieve this goal, one first needs to call genetic variants from NGS data, which requires multiple computationally intensive analysis steps. Unfortunately, there is a lack of an open-source pipeline that can perform all these steps on NGS data in a manner, which is fully automated, efficient, rapid, scalable, modular, user-friendly and fault tolerant. To address this, we introduce xGAP, an extensible Genome Analysis Pipeline, which implements modified GATK best practice to analyze DNA-seq data with the aforementioned functionalities.ResultsxGAP implements massive parallelization of the modified GATK best practice pipeline by splitting a genome into many smaller regions with efficient load-balancing to achieve high scalability. It can process 30× coverage whole-genome sequencing (WGS) data in ∼90 min. In terms of accuracy of discovered variants, xGAP achieves averageF1 scores of 99.37% for single nucleotide variants and 99.20% for insertion/deletions across seven benchmark WGS datasets. We achieve highly consistent results across multiple on-premises (SGE & SLURM) high-performance clusters. Compared to the Churchill pipeline, with similar parallelization, xGAP is 20% faster when analyzing 50× coverage WGS on Amazon Web Service. Finally, xGAP is user-friendly and fault tolerant where it can automatically re-initiate failed processes to minimize required user intervention.Availability and implementationxGAP is available at https://github.com/Adigorla/xgap.Supplementary informationSupplementary data are available atBioinformaticsonline.