A combined long-range phasing and long haplotype imputation method to impute phase for SNP genotypes.

A combined long-range phasing and long haplotype imputation method to impute phase for SNP genotypes.
复制标题

DOI:
10.1186/1297-9686-43-12
复制
发表时间:
2011-03-10
期刊:
Genetics, selection, evolution : GSE
影响因子:
--
通讯作者:
van der Werf JH
van der Werf JH
中科院分区:
其他
文献类型:
--
作者:
Hickey JM;Kinghorn BP;Tier B;Wilson JF;Dunstan N;van der Werf JH

文献摘要

参考文献

被引文献

相似文献

了解标记基因型数据的相位在全基因组关联研究中可能是有用的,因为它使得可以使用通过等位基因的血统或起源的亲本来解释同一性的分析框架,并且它可以通过基因型或序列插补来导致数据量的大幅增加。远程定相和单倍型文库插补构成了一种快速准确的方法来插补SNP数据的相位。开发了一种远程定相和单倍型库插补算法。它结合了来自替代亲本和长单倍型的信息,以不依赖于数据集的家族结构或系谱信息的存在的方式来解析相位。该算法在模拟和真实的牲畜和人类数据集的相位精度和计算效率方面表现良好。在不同大小的模拟和真实的数据集中可以定相的等位基因的百分比通常超过98%,而在模拟数据中不正确定相的等位基因的百分比通常小于0.5%。定相的准确性受数据集大小的影响,数据集大小小于1000的准确性较低,但不受有效人口规模,家庭数据结构,系谱信息的存在或不存在以及SNP密度的影响。该方法计算速度快。与常用的统计方法(fastPHASE)相比,当前方法的定相错误减少了约8%,并且对于小数据集的运行速度快了约26倍。对于较大的数据集,计算时间的差异预计会更大。已经提供了实现这些方法的计算机程序。本研究中开发的算法和软件使大数据集中高密度SNP芯片的常规定相成为可能。
Knowing the phase of marker genotype data can be useful in genome-wide association studies, because it makes it possible to use analysis frameworks that account for identity by descent or parent of origin of alleles and it can lead to a large increase in data quantities via genotype or sequence imputation. Long-range phasing and haplotype library imputation constitute a fast and accurate method to impute phase for SNP data. A long-range phasing and haplotype library imputation algorithm was developed. It combines information from surrogate parents and long haplotypes to resolve phase in a manner that is not dependent on the family structure of a dataset or on the presence of pedigree information. The algorithm performed well in both simulated and real livestock and human datasets in terms of both phasing accuracy and computation efficiency. The percentage of alleles that could be phased in both simulated and real datasets of varying size generally exceeded 98% while the percentage of alleles incorrectly phased in simulated data was generally less than 0.5%. The accuracy of phasing was affected by dataset size, with lower accuracy for dataset sizes less than 1000, but was not affected by effective population size, family data structure, presence or absence of pedigree information, and SNP density. The method was computationally fast. In comparison to a commonly used statistical method (fastPHASE), the current method made about 8% less phasing mistakes and ran about 26 times faster for a small dataset. For larger datasets, the differences in computational time are expected to be even greater. A computer program implementing these methods has been made available. The algorithm and software developed in this study make feasible the routine phasing of high-density SNP chips in large datasets.
DOI: 10.1534/genetics.108.100289
发表时间: 2009-05-01
期刊: GENETICS
影响因子: 3.3
作者:
Habier, D.;Fernando, R. L.;Dekkers, J. C. M.
通讯作者: Dekkers, J. C. M.
DOI: 10.3168/jds.2009-2849
发表时间: 2010-05-01
影响因子: 3.5
作者:
Weigel, K. A.;Van Tassell, C. P.;Wiggans, G. R.
通讯作者: Wiggans, G. R.
DOI: 10.3168/jds.2010-3501
发表时间: 2010-11-01
影响因子: 3.5
作者:
Zhang, Z.;Druet, T.
通讯作者: Druet, T.
DOI: 10.1038/ng.216
发表时间: 2008-09
期刊: Nature genetics
影响因子: 30.8
作者:
Kong A;Masson G;Frigge ML;Gylfason A;Zusmanovich P;Thorleifsson G;Olason PI;Ingason A;Steinberg S;Rafnar T;Sulem P;Mouy M;Jonsson F;Thorsteinsdottir U;Gudbjartsson DF;Stefansson H;Stefansson K
通讯作者: Stefansson K
DOI: 10.1038/nature08625
发表时间: 2009-12-17
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --