CNVkit: Genome-Wide Copy Number Detection and Visualization from Targeted DNA Sequencing.

CNVkit: Genome-Wide Copy Number Detection and Visualization from Targeted DNA Sequencing.
复制标题

DOI:
10.1371/journal.pcbi.1004873
复制
发表时间:
2016-04
影响因子:
4.3
通讯作者:
Bastian BC
Bastian BC
中科院分区:
生物学2区
文献类型:
--
作者:
Talevich E;Shain AH;Botton T;Bastian BC

文献摘要

被引文献

相似文献

生殖系拷贝数变异(CNV)和体细胞拷贝数改变(SCNA)在综合征和癌症中具有重要意义。大规模平行测序越来越多地用于从测序数据中的读取深度的变化推断拷贝数信息。然而,这种方法在靶向重测序的情况下具有局限性,其在选择用于富集的区域之间留下覆盖范围的间隙,并引入与靶捕获和文库制备的效率相关的偏差。我们提出了一种在软件包CNVkit中实现的拷贝数检测方法,该方法使用靶向读段和非特异性捕获的脱靶读段来均匀地推断整个基因组的拷贝数。这种组合既实现了目标区域的外显子水平分辨率,又实现了更大的内含子和基因间区域的足够分辨率,以识别拷贝数变化。特别是,我们成功地推断出拷贝数相当于100个酶的分辨率全基因组从一个平台靶向少至293个基因。在将读数计数标准化为合并参考后,我们评估并校正了解释测序读数深度中大部分无关变异性的三种偏倚来源:GC含量、靶足迹大小和间距以及重复序列。我们将CNVkit的性能与通过阵列比较基因组杂交鉴定的拷贝数变化进行了比较。我们打包了CNVkit的组件,使其易于使用,并提供可视化、重要功能的详细报告以及用于集成到现有分析管道中的导出选项。CNVkit可从https://github.com/etal/cnvkit免费获得。
Germline copy number variants (CNVs) and somatic copy number alterations (SCNAs) are of significant importance in syndromic conditions and cancer. Massively parallel sequencing is increasingly used to infer copy number information from variations in the read depth in sequencing data. However, this approach has limitations in the case of targeted re-sequencing, which leaves gaps in coverage between the regions chosen for enrichment and introduces biases related to the efficiency of target capture and library preparation. We present a method for copy number detection, implemented in the software package CNVkit, that uses both the targeted reads and the nonspecifically captured off-target reads to infer copy number evenly across the genome. This combination achieves both exon-level resolution in targeted regions and sufficient resolution in the larger intronic and intergenic regions to identify copy number changes. In particular, we successfully inferred copy number at equivalent to 100-kilobase resolution genome-wide from a platform targeting as few as 293 genes. After normalizing read counts to a pooled reference, we evaluated and corrected for three sources of bias that explain most of the extraneous variability in the sequencing read depth: GC content, target footprint size and spacing, and repetitive sequences. We compared the performance of CNVkit to copy number changes identified by array comparative genomic hybridization. We packaged the components of CNVkit so that it is straightforward to use and provides visualizations, detailed reporting of significant features, and export options for integration into existing analysis pipelines. CNVkit is freely available from https://github.com/etal/cnvkit.