Phasing of Many Thousands of Genotyped Samples

Phasing of Many Thousands of Genotyped Samples
复制标题

DOI:
10.1016/j.ajhg.2012.06.013
复制
发表时间:
2012-08-10
影响因子:
9.8
通讯作者:
Reich, David
Reich, David
中科院分区:
生物学1区
文献类型:
--
作者:
Wiliams, Amy L.;Patterson, Nick;Reich, David

文献摘要

被引文献

相似文献

单倍型是人类遗传学中大量应用的重要资源,但计算推断的单倍型容易出现转换错误,从而降低其实用性。计算推断单倍型的准确性随着样本大小的增加而增加,尽管正在生成越来越大的基因型数据集,但现有方法需要大量计算资源的事实限制了它们对包含数万或数十万样本的数据集的适用性。在这里,我们提出了 HAPI-UR(无关样本的单倍型推断),这是一种旨在处理不相关和/或三重和二重家族数据的算法,其精度与现有方法相当或更高,并且计算效率高,可应用于 100,000 个或更多样本。我们使用 HAPI-UR 对包含 58,207 个样本的数据集进行阶段化,并表明它实现了实际的运行时间,并且即使使用来自多个种族的样本,开关错误也会随着样本大小而减少。使用包含 16,353 个样本的数据集,我们将 HAPI-UR 与 Beagle、MaCH、IMPUTE2 和 SHAPEIT 进行比较,结果表明 HAPI-UR 的运行速度比所有方法快 18 倍,并且比除 Beagle 之外的其他方法具有更低的切换错误率;通过使用共识定相,运行 HAPI-UR 3 次的切换错误率比 Beagle 稍低,并且速度快了六倍多。我们在另一个具有更高标记密度的数据集上展示了与 Beagle 类似的结果。最后,我们表明 HAPI-UR 比 Beagle 具有更好的运行时扩展属性,因此对于更大的数据集,HAPI-UR 将是实用的,并且将具有更大的运行时优势。 HAPI-UR 可在线获取(请参阅网络资源)。
Haplotypes are an important resource for a large number of applications in human genetics, but computationally inferred haplotypes are subject to switch errors that decrease their utility. The accuracy of computationally inferred haplotypes increases with sample size, and although ever larger genotypic data sets are being generated, the fact that existing methods require substantial computational resources limits their applicability to data sets containing tens or hundreds of thousands of samples. Here, we present HAPI-UR (haplotype inference for unrelated samples), an algorithm that is designed to handle unrelated and/or trio and duo family data, that has accuracy comparable to or greater than existing methods, and that is computationally efficient and can be applied to 100,000 samples or more. We use HAPI-UR to phase a data set with 58,207 samples and show that it achieves practical runtime and that switch errors decrease with sample size even with the use of samples from multiple ethnicities. Using a data set with 16,353 samples, we compare HAPI-UR to Beagle, MaCH, IMPUTE2, and SHAPEIT and show that HAPI-UR runs 18x faster than all methods and has a lower switch-error rate than do other methods except for Beagle; with the use of consensus phasing, running HAPI-UR three times gives a slightly lower switch-error rate than Beagle does and is more than six times faster. We demonstrate results similar to those from Beagle on another data set with a higher marker density. Lastly, we show that HAPI-UR has better runtime scaling properties than does Beagle so that for larger data sets, HAPI-UR will be practical and will have an even larger runtime advantage. HAPI-UR is available online (see Web Resources).