Inferring combined CNV/SNP haplotypes from genotype data

Inferring combined CNV/SNP haplotypes from genotype data
复制标题

DOI:
10.1093/bioinformatics/btq157
复制
发表时间:
2010-06-01
期刊:
影响因子:
5.8
通讯作者:
Coin, Lachlan J. M.
Coin, Lachlan J. M.
中科院分区:
生物学3区
文献类型:
--
作者:
Su, Shu-Yi;Asher, Julian E.;Coin, Lachlan J. M.

文献摘要

被引文献

相似文献

动机:拷贝数变异(拷贝数变异)越来越被认为是个体遗传变异的重要来源,因此研究拷贝数变异的进化史及其对复杂疾病易感性的影响越来越有兴趣。CNV/SNP单倍型对该研究至关重要,但尽管已经提出了许多推断整数拷贝数的方法,但很少有设计用于推断CNV单倍型阶段的方法,而且这些方法都不适用于全基因组规模。在这里,我们提出了一种推断缺失CNV基因型、预测CNV等位基因配置以及从SNP/CNV基因型数据推断CNV单倍型期的方法。我们的方法在polyHap v2.0软件中实现,基于隐马尔可夫模型,该模型模拟了cnv和snp之间的联合单倍型结构。因此,CNVs和SNPs的单倍型期是同时推断的。采用抽样算法来获得每个估计的置信度/可信度。结果:通过对雄性X染色体CNV-SNP单倍型进行配对,获得了二倍体已知期CNV-SNP基因型数据集。我们发现polyHap提供了这些数据集上缺失的CNV基因型、等位基因配置和CNV单倍型期的准确估计。我们将我们的方法应用于非模拟数据集-染色体2上包含短缺失的区域。结果证实,polyHap的准确性可以扩展到现实生活中的数据集。
Motivation: Copy number variations (CNVs) are increasingly recognized as an substantial source of individual genetic variation, and hence there is a growing interest in investigating the evolutionary history of CNVs as well as their impact on complex disease susceptibility. CNV/SNP haplotypes are critical for this research, but although many methods have been proposed for inferring integer copy number, few have been designed for inferring CNV haplotypic phase and none of these are applicable at genome-wide scale. Here, we present a method for inferring missing CNV genotypes, predicting CNV allelic configuration and for inferring CNV haplotypic phase from SNP/CNV genotype data. Our method, implemented in the software polyHap v2.0, is based on a hidden Markov model, which models the joint haplotype structure between CNVs and SNPs. Thus, haplotypic phase of CNVs and SNPs are inferred simultaneously. A sampling algorithm is employed to obtain a measure of confidence/credibility of each estimate.Results: We generated diploid phase-known CNV-SNP genotype datasets by pairing male X chromosome CNV-SNP haplotypes. We show that polyHap provides accurate estimates of missing CNV genotypes, allelic configuration and CNV haplotypic phase on these datasets. We applied our method to a non-simulated dataset-a region on Chromosome 2 encompassing a short deletion. The results confirm that polyHap's accuracy extends to real-life datasets.