A One-Penny Imputed Genome from Next-Generation Reference Panels

A One-Penny Imputed Genome from Next-Generation Reference Panels
复制标题

DOI:
10.1016/j.ajhg.2018.07.015
复制
发表时间:
2018-09-06
影响因子:
9.8
通讯作者:
Browning, Sharon R.
Browning, Sharon R.
中科院分区:
生物学1区
文献类型:
--
作者:
Browning, Brian L.;Zhou, Ying;Browning, Sharon R.

文献摘要

被引文献

相似文献

基因型插补通常在全基因组关联研究中进行,因为它大大增加了可以测试与性状关联的标记的数量。一般而言,应使用可用的最大参考样本组进行基因型插补,因为准确插补的变异数量随参考样本组大小而增加。然而,使用较大参考样本组的一个障碍是增加了插补的计算成本。我们提出了一种新的基因型插补方法,Beagle 5.0,这大大降低了从大型参考面板插补的计算成本。我们使用1000个基因组项目数据、单倍型参考联盟数据以及10 k、100 k、1 M和10 M参考样本的模拟数据比较了Beagle 5.0与Beagle 4.1、Impute 4、Minimac 3和Minimac 4。所有方法产生几乎相同的准确度,但Beagle 5.0具有最低的计算时间和最佳的计算时间缩放与参考面板大小的增加。对于10 k、100 k、1 M和10 M参考样本以及1,000个定相目标样本,Beagle 5. 0的计算时间比最快的替代方法快33(10 k)、123(100 k)、433(1 M)和5333(10 M)。来自Amazon Elastic Compute Cloud的成本数据显示,Beagle 5.0可以从1,000万个参考样本到1,000个分阶段目标样本进行全基因组插补,每个样本的成本不到1美分。
Genotype imputation is commonly performed in genome-wide association studies because it greatly increases the number of markers that can be tested for association with a trait. In general, one should perform genotype imputation using the largest reference panel that is available because the number of accurately imputed variants increases with reference panel size. However, one impediment to using larger reference panels is the increased computational cost of imputation. We present a new genotype imputation method, Beagle 5.0, which greatly reduces the computational cost of imputation from large reference panels. We compare Beagle 5.0 with Beagle 4.1, Impute4, Minimac3, and Minimac4 using 1000 Genomes Project data, Haplotype Reference Consortium data, and simulated data for 10k, 100k, 1M, and 10M reference samples. All methods produce nearly identical accuracy, but Beagle 5.0 has the lowest computation time and the best scaling of computation time with increasing reference panel size. For 10k, 100k, 1M, and 10M reference samples and 1,000 phased target samples, Beagle 5.0' s computation time is 33 (10k), 123 (100k), 433 (1M), and 5333 (10M) faster than the fastest alternative method. Cost data from the Amazon Elastic Compute Cloud show that Beagle 5.0 can perform genome-wide imputation from 10M reference samples into 1,000 phased target samples at a cost of less than one US cent per sample.