A fast and flexible statistical model for large-scale population genotype data: Applications to inferring missing genotypes and haplotypic phase

A fast and flexible statistical model for large-scale population genotype data: Applications to inferring missing genotypes and haplotypic phase
复制标题

DOI:
10.1086/502802
复制
发表时间:
2006-04-01
影响因子:
9.8
通讯作者:
Stephens, M
Stephens, M
中科院分区:
生物学1区
文献类型:
--
作者:
Scheet, P;Stephens, M

文献摘要

被引文献

相似文献

我们提出了一个统计模型的遗传变异模式的样本无关的个人从自然种群。该模型基于这样的想法,即在短区域内,群体中的单倍型倾向于聚集成相似单倍型的组。为了捕捉这样一个事实,即,由于重组,这种聚类往往是本地的性质,我们的模型允许集群成员连续沿着染色体根据隐马尔可夫模型变化。这种方法是灵活的,允许“块状”模式的连锁不平衡(LD)和LD随着距离的逐渐下降。由此产生的模型也是快速的,因此,对于大数据集是可行的。例如,在一个实施例中,数以千计的个体在数十万个标记处键入)。我们说明了该模型的实用性,将其应用到密集的单核苷酸多态性基因型数据的任务,填补缺失的基因型和估计单倍型阶段。对于缺失基因型的插补,基于该模型的方法与现有方法一样准确或更准确。对于单倍型估计,点估计的准确性略低于现有的最佳方法(例如。例如,在一个实施例中,对于来自HapMap项目的无关人类多态性研究中心的个体,我们的方法的转换误差为0.055,PHASE的转换误差为0.051),但需要一小部分计算成本。此外,我们证明,该模型准确地反映了其估计的不确定性,在使用该模型计算的概率近似以及校准。本文中描述的方法在软件包fastPHASE中实现,该软件包可从Stephens Lab网站获得。
We present a statistical model for patterns of genetic variation in samples of unrelated individuals from natural populations. This model is based on the idea that, over short regions, haplotypes in a population tend to cluster into groups of similar haplotypes. To capture the fact that, because of recombination, this clustering tends to be local in nature, our model allows cluster memberships to change continuously along the chromosome according to a hidden Markov model. This approach is flexible, allowing for both "block-like" patterns of linkage disequilibrium ( LD) and gradual decline in LD with distance. The resulting model is also fast and, as a result, is practicable for large data sets ( e. g., thousands of individuals typed at hundreds of thousands of markers). We illustrate the utility of the model by applying it to dense single-nucleotide-polymorphism genotype data for the tasks of imputing missing genotypes and estimating haplotypic phase. For imputing missing genotypes, methods based on this model are as accurate or more accurate than existing methods. For haplotype estimation, the point estimates are slightly less accurate than those from the best existing methods ( e. g., for unrelated Centre d'Etude du Polymorphisme Humain individuals from the HapMap project, switch error was 0.055 for our method vs. 0.051 for PHASE) but require a small fraction of the computational cost. In addition, we demonstrate that the model accurately reflects uncertainty in its estimates, in that probabilities computed using the model are approximately well calibrated. The methods described in this article are implemented in a software package, fastPHASE, which is available from the Stephens Lab Web site.