Shall genomic correlation structure be considered in copy number variants detection?

Shall genomic correlation structure be considered in copy number variants detection?
复制标题

拷贝数变异检测中是否应考虑基因组相关结构?

DOI:
10.1093/bib/bbab215
复制
发表时间:
2021
影响因子:
9.5
通讯作者:
Xiao,Feifei
Xiao,Feifei
中科院分区:
生物学2区
文献类型:
--
作者:
Qin,Fei;Luo,Xizhi;Cai,Guoshuai;Xiao,Feifei

文献摘要

相似文献

拷贝数变异已被确定为与疾病易感性相关的基因组变异的主要来源。随着全外显子组测序(WES)技术的出现,已经产生了大量的WES数据,允许识别蛋白质编码区的拷贝数变异(CNV),并进行直接的功能解释。我们先前已经显示了阵列数据中基因组相关结构的证据,并开发了一种新的染色体断点检测算法LDcnv,该算法通过以系统建模的方式整合相关结构来显着提高检测能力。然而,WES数据中是否存在基因组相关性以及这种相关性结构整合如何提高CNV检测准确性仍有待探索。在这项研究中,我们首先探索了WES数据的相关结构,使用1000个基因组计划的数据。真实的原始读取深度和中值归一化数据均显示出相关性结构的强有力证据。出于这一事实,我们提出了一个基于相关性的方法,CORRseq,作为一个新的版本的LDcnv算法在分析WES数据。在广泛的模拟研究和来自1000个基因组计划的真实的数据分析中评估了CORRseq的性能。CORRseq在检测中等和大CNV方面优于现有方法。总之,在检测相对长的CNV时,对基因组相关结构进行建模将是更有利的。该研究为利用NGS数据进行CNV检测的方法学发展提供了重要见解。
Copy number variation has been identified as a major source of genomic variation associated with disease susceptibility. With the advent of whole-exome sequencing (WES) technology, massive WES data have been generated, allowing for the identification of copy number variants (CNVs) in the protein-coding regions with direct functional interpretation. We have previously shown evidence of the genomic correlation structure in array data and developed a novel chromosomal breakpoint detection algorithm, LDcnv, which showed significantly improved detection power through integrating the correlation structure in a systematic modeling manner. However, it remains unexplored whether the genomic correlation exists in WES data and how such correlation structure integration can improve the CNV detection accuracy. In this study, we first explored the correlation structure of the WES data using the 1000 Genomes Project data. Both real raw read depth and median-normalized data showed strong evidence of the correlation structure. Motivated by this fact, we proposed a correlation-based method, CORRseq, as a novel release of the LDcnv algorithm in profiling WES data. The performance of CORRseq was evaluated in extensive simulation studies and real data analysis from the 1000 Genomes Project. CORRseq outperformed the existing methods in detecting medium and large CNVs. In conclusion, it would be more advantageous to model genomic correlation structure in detecting relatively long CNVs. This study provides great insights for methodology development of CNV detection with NGS data.