ParseCNV2: a versatile and integrated tool for copy number variation association studies.

ParseCNV2: a versatile and integrated tool for copy number variation association studies.
复制标题

ParseCNV2:用于拷贝数变异关联研究的多功能集成工具。

DOI:
10.1038/s41431-022-01280-x
复制
发表时间:
2023
期刊:
European journal of human genetics : EJHG
影响因子:
--
通讯作者:
Sanna-Cherchi,Simone
Sanna-Cherchi,Simone
中科院分区:
--
文献类型:
--
作者:
Lim,TzeY;Verbitsky,Miguel;Sanna-Cherchi,Simone

文献摘要

相似文献

基因组结构变异(SV),包括拷贝数变异(CNV)和拷贝中性变异,是研究不足的基因组变异类型[1,2]。它们可以包含大的DNA片段,并且包括多个基因和调控元件,并且与参考基因组相比,它们以可变的结构或拷贝存在。相比之下,在基于家族或病例对照关联研究中发现的常见或罕见的单核苷酸变异(SNV)一直是绝大多数人类遗传学研究的对象,但单独无法完全解释复杂疾病的遗传基础[3]。CNVs在人类疾病易感性中的作用已经得到了很好的证实,但它们对这种“缺失的遗传性”的全部贡献仍然在很大程度上未知。与SNV相比,SV/CNV研究较少的两个主要原因是CNV区域(CNVr)的准确基因分型困难,以及有效进行充分动力,无假设的病例对照关联研究的挑战,从而限制了新的发现。在过去的世纪中,结构变异鉴定和分析领域已经从分辨率有限的非常简单和有针对性的方法发展到最近的基于阵列或基于测序的全基因组方法(图1A)。因此,使用基于克隆的比较基因组杂交(aCGH),高密度SNP基因分型阵列和最近的人类基因组深度测序来进行SV/CNV分析,我们已经实现了更高的分辨率,但也增加了数据分析的显着复杂性。现代大规模病例对照研究通常需要汇集来自不同队列的数据,这些数据可能使用不同的技术(即DNA微阵列,外显子组或基因组测序),相同技术的不同捕获版本,不同的处理程序以及不同的CNV调用者生成。因此,整合这些不同的数据集,解释队列异质性,批次效应和数据兼容性问题的荟萃分析提出了重大挑战。因此,关联分析和解释方法的不断发展是充分理解结构变异对健康和疾病的贡献的关键。在本期中,Glessner等人报道了ParseCNV 2 [4],这是一种新的软件更新,它执行CNV关联,具有额外的功能,可以原生地支持基因分型阵列和测序数据。ParseCNV 2中的关键增加包括:1)代码修改,以开发统一的CNV变异调用格式(VCF)解析器,因为当前的VCF文件格式缺乏报告CNV的标准约定; 2)新选项,以在调整时对定量和二元性状执行线性或逻辑回归测试。
Genomic structural variants (SVs), including copy number variations (CNVs) and copy neutral variation, are understudied classes of genomic variations [1, 2]. They can comprise large DNA segments and include multiple genes and regulatory element, and they present in variable structure or copies in comparison to the reference genome. In contrast, single nucleotide variants (SNV), common or rare, identified in family-based or case-control association studies, have been the object of the vast majority of human genetic studies, but alone cannot fully explain the genetic basis of complex diseases [3]. The role of CNVs in human disease predisposition is well established, but their full contribution to this “missing heritability” still remains largely unknown. Two of the main reasons for the paucity of studies on SV/CNVs as compared to SNVs are in the difficulty of accurately genotyping CNV regions (CNVr) and in the challenges to effectively conduct an adequately powered, hypothesis-free case-control association study, thus limiting new discoveries. In the past century the field of structural variation identification and analysis has evolved from very simple and targeted approaches with limited resolution to more recent array-based or sequencing-based genome-wide approaches (Fig. 1 A). Hence, using clone-based comparative genomic hydbridization (aCGH), high density SNP genotyping array and more recent deep sequencing of the human genome to conduct SV/CNV analysis, we have achieved a far higher resolution but also added significant complexity in data analysis. Modern large-scale case-control studies often require pooling data from different cohorts which may be generated using different technologies (ie DNA microarrays, exome or genome sequencing), different capture versions of the same technology, different processing procedures, as well as diverse CNV callers. Therefore, meta-analyses that integrate such diverse datasets accounting for cohort heterogeneity, batch effects, and data compatibility issues present significant challenges. Continuous evolution of methods for association analysis and interpretation is therefore key to fully understand the contribution of structural variants to health and disease.In this issue, Glessner et al, report on ParseCNV2 [4], a new update of a software that performs CNV association with extra functionalities to natively support genotyping array and sequencing data. The critical additions in ParseCNV2 include: 1) a code revamp to develop a unified CNV variant call format (VCF) parser, as current VCF file format lacks a standard convention for reporting CNVs; 2) new options to perform either linear or logistic regression tests for quantitative and binary traits while adjusting