PennCNV: An integrated hidden Markov model designed for high-resolution copy number variation detection in whole-genome SNP genotyping data

PennCNV: An integrated hidden Markov model designed for high-resolution copy number variation detection in whole-genome SNP genotyping data
复制标题

DOI:
10.1101/gr.6861907
复制
发表时间:
2007-11-01
期刊:
影响因子:
7
通讯作者:
Bucan, Maja
Bucan, Maja
中科院分区:
生物学1区
文献类型:
--
作者:
Wang, Kai;Li, Mingyao;Bucan, Maja

文献摘要

被引文献

相似文献

需要对拷贝数变异 (CNV) 进行全面的识别和编目,以提供人类遗传变异的完整视图。之前的实验设计中CNV检测的分辨率仅限于数十或数百千碱基。在这里,我们介绍 PennCNV,一种基于隐马尔可夫模型 (HMM) 的方法,用于从 Illumina 高密度 SNP 基因分型数据中以千碱基分辨率检测 CNV。该算法结合了多个信息源,包括每个 SNP 标记的总信号强度和等位基因强度比、相邻 SNP 之间的距离、SNP 的等位基因频率以及可用的谱系信息。我们应用 PennCNV 对 112 个 HapMap 个体生成的基因分型数据进行分析;平均而言,我们为每个个体检测到 -27 个 CNV,中位大小为 -12 kb。排除类淋巴母细胞系中常见的重排,在父母中未检测到的后代 CNV 比例 (CNV-NDP) 为 3.3%。我们的结果证明了通过高密度 SNP 基因分型对 CNV 进行全基因组精细定位的可行性。
Comprehensive identification and cataloging of copy number variations (CNVs) is required to provide a complete view of human genetic variation. The resolution of CNV detection in previous experimental designs has been limited to tens or hundreds of kilobases. Here we present PennCNV, a hidden Markov model (HMM) based approach, for kilobase-resolution detection of CNVs from Illumina high-density SNP genotyping data. This algorithm incorporates multiple sources of information, including total signal intensity and allelic intensity ratio at each SNP marker, the distance between neighboring SNPs, the allele frequency of SNPs, and the pedigree information where available. We applied PennCNV to genotyping data generated for 112 HapMap individuals; on average, we detected -27 CNVs for each individual with a median size of -12 kb. Excluding common rearrangements in lymphoblastoid cell lines, the fraction of CNVs in offspring not detected in parents (CNV-NDPs) was 3.3%. Our results demonstrate the feasibility of whole-genome fine-mapping of CNVs via high-density SNP genotyping.