ParseCNV2: a versatile and integrated tool for copy number variation association studies.
ParseCNV2: a versatile and integrated tool for copy number variation association studies.
复制标题
ParseCNV2:用于拷贝数变异关联研究的多功能集成工具。
DOI:
10.1038/s41431-022-01280-x
复制
发表时间:
2023
期刊:
影响因子:
--
通讯作者:
Sanna-Cherchi,Simone
中科院分区:
文献类型:
--
作者:
Lim,TzeY;Verbitsky,Miguel;Sanna-Cherchi,Simone
Genomic structural variants (SVs), including copy number variations (CNVs) and copy neutral variation, are understudied classes of genomic variations [1, 2]. They can comprise large DNA segments and include multiple genes and regulatory element, and they present in variable structure or copies in comparison to the reference genome. In contrast, single nucleotide variants (SNV), common or rare, identified in family-based or case-control association studies, have been the object of the vast majority of human genetic studies, but alone cannot fully explain the genetic basis of complex diseases [3]. The role of CNVs in human disease predisposition is well established, but their full contribution to this “missing heritability” still remains largely unknown. Two of the main reasons for the paucity of studies on SV/CNVs as compared to SNVs are in the difficulty of accurately genotyping CNV regions (CNVr) and in the challenges to effectively conduct an adequately powered, hypothesis-free case-control association study, thus limiting new discoveries. In the past century the field of structural variation identification and analysis has evolved from very simple and targeted approaches with limited resolution to more recent array-based or sequencing-based genome-wide approaches (Fig. 1 A). Hence, using clone-based comparative genomic hydbridization (aCGH), high density SNP genotyping array and more recent deep sequencing of the human genome to conduct SV/CNV analysis, we have achieved a far higher resolution but also added significant complexity in data analysis. Modern large-scale case-control studies often require pooling data from different cohorts which may be generated using different technologies (ie DNA microarrays, exome or genome sequencing), different capture versions of the same technology, different processing procedures, as well as diverse CNV callers. Therefore, meta-analyses that integrate such diverse datasets accounting for cohort heterogeneity, batch effects, and data compatibility issues present significant challenges. Continuous evolution of methods for association analysis and interpretation is therefore key to fully understand the contribution of structural variants to health and disease.In this issue, Glessner et al, report on ParseCNV2 [4], a new update of a software that performs CNV association with extra functionalities to natively support genotyping array and sequencing data. The critical additions in ParseCNV2 include: 1) a code revamp to develop a unified CNV variant call format (VCF) parser, as current VCF file format lacks a standard convention for reporting CNVs; 2) new options to perform either linear or logistic regression tests for quantitative and binary traits while adjusting