The Use of Family Relationships and Linkage Disequilibrium to Impute Phase and Missing Genotypes in Up to Whole-Genome Sequence Density Genotypic Data

The Use of Family Relationships and Linkage Disequilibrium to Impute Phase and Missing Genotypes in Up to Whole-Genome Sequence Density Genotypic Data
复制标题

DOI:
10.1534/genetics.110.113936
复制
发表时间:
2010-08-01
期刊:
影响因子:
3.3
通讯作者:
Goddard, Mike
Goddard, Mike
中科院分区:
生物学2区
文献类型:
--
作者:
Meuwissen, Theo;Goddard, Mike

文献摘要

被引文献

相似文献

提出了一种用于相位和缺失基因型插入的新方法——连锁不平衡多位点迭代剥离(LDMIP)。LDMIP对每个基因座执行迭代剥离步骤,这说明了家族数据,并使用前向向后算法在基因座之间积累信息。利用单倍型对之间的标记相似性来推测可能缺失的基因型和期,这依赖于紧密连锁标记之间的连锁不平衡。在这一步后,再次应用迭代剥离/向前-向后组合算法,直到收敛。每次迭代的计算与谱系中的标记数和个体数呈线性关系,这使得LDMIP非常适合大量标记和/或大量个体。每次迭代计算的规模与等位基因的数量成二次关系,这意味着双等位基因标记是首选的。在随机缺失基因型高达15%的情况下,缺失基因型的输入错误率为99%。
A novel method, called linkage disequilibrium multilocus iterative peeling (LDMIP), for the imputation of phase and missing genotypes is developed. LDMIP performs an iterative peeling step for every locus, which accounts for the family data, and uses a forward-backward algorithm to accumulate information across loci. Marker similarity between haplotype pairs is used to impute possible missing genotypes and phases, which relies on the linkage disequilibrium between closely linked markers. After this imputation step, the combined iterative peeling/forward-backward algorithm is applied again, until convergence. The calculations per iteration scale linearly with number of markers and number of individuals in the pedigree, which makes LDMIP well suited to large numbers of markers and/or large numbers of individuals. Per iteration calculations scale quadratically with the number of alleles, which implies biallelic markers are preferred. In a situation with up to 15% randomly missing genotypes, the error rate of the imputed genotypes was 99% of missing genotypes are imputed correctly.