GIGI: An Approach to Effective Imputation of Dense Genotypes on Large Pedigrees

GIGI: An Approach to Effective Imputation of Dense Genotypes on Large Pedigrees
复制标题

DOI:
10.1016/j.ajhg.2013.02.011
复制
发表时间:
2013-04-04
影响因子:
9.8
通讯作者:
Wijsman, Ellen M.
Wijsman, Ellen M.
中科院分区:
生物学1区
文献类型:
--
作者:
Cheung, Charles Y. K.;Thompson, Elizabeth A.;Wijsman, Ellen M.

文献摘要

被引文献

相似文献

最近出现的常见疾病罕见变异假说重新引起了人们对使用大型家系来识别罕见因果变异的兴趣,利用现代测序平台进行基因分型在寻找此类变异中越来越常见,但仍然昂贵,而且通常仅限于每个家系的几个受试者。在以人群为基础的样本中,广泛使用的是基因归类,因此不需要额外的基因分型。我们现在介绍一种类似的方法,它能够在计算上有效地对大型家系进行归类。我们的方法从马尔可夫链蒙特卡罗采样器中采样遗传向量(IV),方法是对稀疏框架标记集的基因类型进行条件处理。缺失的基因类型是根据这些静脉输液以及观察到的在受试者子集上可用的密集基因类型来概率推断的。我们在GIGI程序中实现了我们的方法,并在模拟和真实的大型家系上对该方法进行了评估。对于一个真实的家系,我们还比较了从这种方法获得的推算结果与基于人口的推算程序Beagle的推算结果。我们证明了我们的基于系谱的方法以高准确度归因于许多等位基因。与基于总体的推定相比,这种方法对稀有等位基因的调用更为准确,而且不需要外部参考样本。我们还评估了改变其他参数的影响,包括框架面板的标记类型和密度、调用基因型的阈值和群体等位基因频率。通过利用已经在大型家系上分析的现有基因类型的信息,我们的方法可以促进在追求罕见因果变异的过程中经济有效地使用序列数据。
Recent emergence of the common-disease-rare-variant hypothesis has renewed interest in the use of large pedigrees for identifying rare causal variants Genotyping with modem sequencing platforms is increasingly common in the search for such variants but remains expensive and often is limited to only a few subjects per pedigree. In population-based samples, genotype imputation is widely used so that additional genotyping is not needed. We now introduce an analogous approach that enables computationally efficient imputation in large pedigrees. Our approach samples inheritance vectors (IVs) from a Markov Chain Monte Carlo sampler by conditioning on genotypes from a sparse set of framework markers. Missing genotypes are probabilistically inferred from these IVs along with observed dense genotypes that are available on a subset of subjects. We implemented our approach in the Genotype Imputation Given Inheritance (GIGI) program and evaluated the approach on both simulated and real large pedigrees. With a real pedigree, we also compared imputed results obtained from this approach with those from the population-based imputation program BEAGLE. We demonstrated that our pedigree-based approach imputes many alleles with high accuracy. It is much more accurate for calling rare alleles than is population-based imputation and does not require an outside reference sample. We also evaluated the effect of varying other parameters, including the marker type and density of the framework panel, threshold for calling genotypes, and population allele frequencies. By leveraging information from existing genotypes already assayed on large pedigrees, our approach can facilitate cost-effective use of sequence data in the pursuit of rare causal variants.