Genotype Imputation with Millions of Reference Samples

Genotype Imputation with Millions of Reference Samples
复制标题

DOI:
10.1016/j.ajhg.2015.11.020
复制
发表时间:
2016-01-07
影响因子:
9.8
通讯作者:
Browning, Sharon R.
Browning, Sharon R.
中科院分区:
生物学1区
文献类型:
--
作者:
Browning, Brian L.;Browning, Sharon R.

文献摘要

被引文献

相似文献

我们提出了一个基因型插补方法,规模以百万计的参考样本。基于Li和Stephens模型并在Beagle v.4.1中实现的插补方法是并行化的,并且具有内存效率,使得它非常适合于多核计算机处理器。它通过将概率模型限制在目标样本中基因分型的标记物上,并通过执行线性插值来插补未分型的变异体,从而实现快速、准确和记忆高效的基因型插补。我们使用1000个基因组计划数据,UK10K计划数据和模拟数据比较Beagle v.4.1与Impute 2和Minimac 3。所有三种方法具有相似的精度,但不同的内存需求和不同的计算时间。当输入来自50,000个参考样品的10 Mb序列数据时,Beagle的吞吐量比我们计算机服务器上的Impute 2的吞吐量大100倍以上。当以VCF格式输入来自200,000个参考样品的10 Mb序列数据时,Minimac3每个计算线程消耗的内存是Beagle的26倍,CPU时间是Beagle的15倍。我们证明,比格犬v.4.1规模更大的参考面板进行插补从一个模拟的参考面板有5万个样本和一个标记的平均密度为每四个碱基对的标记。
We present a genotype imputation method that scales to millions of reference samples. The imputation method, based on the Li and Stephens model and implemented in Beagle v.4.1, is parallelized and memory efficient, making it well suited to multi-core computer processors. It achieves fast, accurate, and memory-efficient genotype imputation by restricting the probability model to markers that are genotyped in the target samples and by performing linear interpolation to impute ungenotyped variants. We compare Beagle v.4.1 with Impute2 and Minimac3 by using 1000 Genomes Project data, UK10K Project data, and simulated data. All three methods have similar accuracy but different memory requirements and different computation times. When imputing 10 Mb of sequence data from 50,000 reference samples, Beagle's throughput was more than 100x greater than Impute2's throughput on our computer servers. When imputing 10 Mb of sequence data from 200,000 reference samples in VCF format, Minimac3 consumed 26x more memory per computational thread and 15x more CPU time than Beagle. We demonstrate that Beagle v.4.1 scales to much larger reference panels-by performing imputation from a simulated reference panel having 5 million samples and a mean marker density of one marker per four base pairs.