COPILOT: a Containerised wOrkflow for Processing ILlumina genOtyping daTa

COPILOT: a Containerised wOrkflow for Processing ILlumina genOtyping daTa
复制标题

COPILOT:用于处理 ILlumina 基因分型数据的容器化工作流程

DOI:
10.1101/2021.07.26.453753
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Patel H
Patel H
中科院分区:
--
文献类型:
--
作者:
Patel H

文献摘要

参考文献

相似文献

背景Illumina基因分型微阵列产生图像格式的数据,这些数据由特定平台的软件GenomeStudio处理,然后进行一系列复杂的生物信息学分析。这一过程可能会耗时,导致重复性错误,对于新手生物信息学家来说是一项艰巨的任务。结果在这里,我们介绍了Copilot(处理ILLumina基因分型数据的容器化工作流)协议,它提供了在GenomeStudio中处理原始Illumina基因数据的深入而清晰的指南,随后是一个容器化工作流,用于自动化一系列涉及GWAS质量控制(QC)的复杂生物信息学分析。Copilot方案应用于两个独立的队列,分别由2791个和479个样本组成,这些样本分别在Infinium Global Screen(GSA)阵列和Infinium H3Africa联合阵列上进行基因分型(约750,000个标记)和Infinium H3 Africa联合阵列(约2,200,000个标记)。根据Copilot协议,样本呼叫率的平均样本质量改善了1.24%,低质量样本的改善尤为显著。例如,在处理的3270个样本中,141个样本的初始样本呼叫率低于98%,平均为96.6%(95%可信区间95.6-97.7%),这被认为低于典型GWAS分析的可接受样本呼叫率阈值。然而,按照Copilot方案,所有141个样本在质量控制后的呼叫率都在98%以上,平均为99.6%(95%可信区间99.5-99.7%)。此外,Copilot管道通过祖先估计自动识别潜在的数据问题,包括性别差异、杂合性异常值、相关个体和群体异常值。结论Copilot协议使Illumina基因分型数据的处理透明、轻松和可重复性。该容器可部署在多个平台上,提高了数据质量,最终产品是可供分析的plink格式数据,具有全面和交互的摘要报告,指导用户进行进一步的数据分析。
BackgroundThe Illumina genotyping microarrays generate data in image format, which is processed by the platform-specific software GenomeStudio, followed by an array of complex bioinformatics analyses. This process can be time-consuming, lead to reproducibility errors, and be a daunting task for novice bioinformaticians.ResultsHere we introduce the COPILOT (Containerised wOrkflow forProcessingILlumina genOtyping daTa) protocol, which provides an in-depth and clear guide to process raw Illumina genotype data in GenomeStudio, followed by a containerised workflow to automate an array of complex bioinformatics analyses involved in a GWAS quality control (QC). The COPILOT protocol was applied to two independent cohorts consisting of 2791 and 479 samples genotyped on the Infinium Global Screening (GSA) array with Multi-disease (MD) drop-in (~750,000 markers) and the Infinium H3Africa consortium array (~2,200,000 markers) respectively. Following the COPILOT protocol, an average sample quality improvement of 1.24% was observed across sample call rates, with notable improvement for low-quality samples. For example, from the 3270 samples processed, 141 samples had an initial sample call rate below 98%, averaging 96.6% (95% CI 95.6-97.7%), which is considered below the acceptable sample call rate threshold for a typical GWAS analysis. However, following the COPILOT protocol, all 141 samples had a call rate above 98% after QC and averaged 99.6% (95% CI 99.5-99.7%). In addition, the COPILOT pipeline automatically identified potential data issues, including gender discrepancies, heterozygosity outliers, related individuals, and population outliers through ancestry estimation.ConclusionsThe COPILOT protocol makes processing Illumina genotyping data transparent, effortless and reproducible. The container is deployable on multiple platforms, improves data quality, and the end product is analysis-ready PLINK formatted data, with a comprehensive and interactive summary report to guide the user for further data analyses.
来自1,092个人基因组的遗传变异的综合图。
DOI: 10.1038/nature11632
发表时间: 2012-11-01
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --