Rapid and accurate haplotype phasing and missing-data inference for whole-genome association studies by use of localized haplotype clustering

Rapid and accurate haplotype phasing and missing-data inference for whole-genome association studies by use of localized haplotype clustering
复制标题

DOI:
10.1086/521987
复制
发表时间:
2007-11-01
影响因子:
9.8
通讯作者:
Browning, Brian L.
Browning, Brian L.
中科院分区:
生物学1区
文献类型:
--
作者:
Browning, Sharon R.;Browning, Brian L.

文献摘要

被引文献

相似文献

由于获得了大量的数据,全基因组关联研究提出了许多新的统计和计算挑战。其中一个挑战是单倍型推断;针对候选基因研究中的小数据集设计的单倍型推断方法不能很好地扩展到全基因组关联研究中大量的基因分型个体。我们提出了一种新的方法和软件来推断单倍型阶段和缺失数据,可以准确地从全基因组关联研究中获得阶段数据,并且我们首次比较了真实和模拟数据集上的单倍型推断方法与数千个基因分型个体的单倍型推断方法。我们发现,对于包含数千个个体和密集遗传标记的大型数据集,我们的方法在速度和准确性方面都优于现有的方法,并使用我们的方法在3.1天的计算时间内对3,002个个体的490,032个标记进行了基因分型,99%的掩蔽等位基因被正确推定。我们的方法是在Beagle软件包中实现的,该软件包是免费提供的。
Whole-genome association studies present many new statistical and computational challenges due to the large quantity of data obtained. One of these challenges is haplotype inference; methods for haplotype inference designed for small data sets from candidate-gene studies do not scale well to the large number of individuals genotyped in whole-genome association studies. We present a new method and software for inference of haplotype phase and missing data that can accurately phase data from whole-genome association studies, and we present the first comparison of haplotype-inference methods for real and simulated data sets with thousands of genotyped individuals. We find that our method outperforms existing methods in terms of both speed and accuracy for large data sets with thousands of individuals and densely spaced genetic markers, and we use our method to phase a real data set of 3,002 individuals genotyped for 490,032 markers in 3.1 days of computing time, with 99% of masked alleles imputed correctly. Our method is implemented in the Beagle software package, which is freely available.