ParseCNV2: efficient sequencing tool for copy number variation genome-wide association studies

ParseCNV2: efficient sequencing tool for copy number variation genome-wide association studies
复制标题

DOI:
10.1038/s41431-022-01222-7
复制
发表时间:
2022-11-01
影响因子:
5.2
通讯作者:
Hakonarson, Hakon
Hakonarson, Hakon
中科院分区:
生物学2区
文献类型:
--
作者:
Glessner, Joseph T.;Li, Jin;Hakonarson, Hakon

文献摘要

被引文献

相似文献

改进的拷贝数变异(CNV)检测仍然是算法开发的重点领域;然而,CNV治疗和疾病关联方法仍处于起步阶段。目前专注于候选CNVs的做法是,研究人员研究他们认为是病理性的特定CNVs,同时丢弃其他CNVs,避免在无假设的GWAS中考虑CNVs的全谱。为了解决这个问题,我们提出了一种新一代的CNV关联方法,通过原生支持流行的VCF规范,用于测序衍生的变体以及使用PennCNV格式的SNP阵列调用。该代码快速高效,允许分析大型(> 100,000个样本)队列,而无需在计算集群上划分数据。这些脚本被压缩到一个单一的工具中,以促进简单性和最佳实践。CNV治疗前和后协会是严格支持和强调,以产生最高质量的可靠结果。我们对两个大型数据集进行了基准测试,包括UK Biobank(n> 450,000)和CAG Biobank(n > 350,000),这两个数据集都是在> 0.5M探针下进行基因分型的,用于我们的输入文件。ParseCNV自2008年以来一直得到积极支持和开发。ParseCNV 2为将CNV关联正式化提供了一个重要的补充,以便与GWAS目录中的SNP关联一起纳入。临床CNV优先化、交互式质量控制(QC)和协变量调整是ParseCNV 2与ParseCNV相比的革命性新功能。该软件可在https://github.com/CAG-CNV/ParseCNV2上免费获得。
Improved copy number variation (CNV) detection remains an area of heavy emphasis for algorithm development; however, both CNV curation and disease association approaches remain in its infancy. The current practice of focusing on candidate CNVs, where researchers study specific CNVs they believe to be pathological while discarding others, refrains from considering the full spectrum of CNVs in a hypothesis-free GWAS. To address this, we present a next-generation approach to CNV association by natively supporting the popular VCF specification for sequencing-derived variants as well as SNP array calls using a PennCNV format. The code is fast and efficient, allowing for the analysis of large (>100,000 sample) cohorts without dividing up the data on a compute cluster. The scripts are condensed into a single tool to promote simplicity and best practices. CNV curation pre and post-association is rigorously supported and emphasized to yield reliable results of highest quality. We benchmarked two large datasets, including the UK Biobank (n> 450,000) and CAG Biobank (n > 350,000) both of which are genotyped at >0.5 M probes, for our input files. ParseCNV has been actively supported and developed since 2008. ParseCNV2 presents a critical addition to formalizing CNV association for inclusion with SNP associations in GWAS Catalog. Clinical CNV prioritization, interactive quality control (QC), and adjustment for covariates are revolutionary new features of ParseCNV2 vs. ParseCNV. The software is freely available at https://github.com/CAG-CNV/ParseCNV2.