Fast individual ancestry inference from DNA sequence data leveraging allele frequencies for multiple populations.

Fast individual ancestry inference from DNA sequence data leveraging allele frequencies for multiple populations.
复制标题

DOI:
10.1186/s12859-014-0418-7
复制
发表时间:
2015-01-16
期刊:
影响因子:
3
通讯作者:
Libiger O
Libiger O
中科院分区:
生物学4区
文献类型:
--
作者:
Bansal V;Libiger O

文献摘要

参考文献

被引文献

相似文献

从遗传数据中估计个体祖先对疾病关联研究的分析、了解人类种群历史和解释个人基因组变异是有用的。需要新的计算效率高的方法来进行祖先推断,这些方法可以有效地利用与不同人群相关的等位基因频率的现有信息,并且可以直接与DNA序列读取一起工作。我们描述了一种快速估计已知参考种群对个体遗传祖先的相对贡献的方法。我们的方法利用参考群体和个体基因型或序列数据中的等位基因频率,使用BFGS优化算法获得全局混合比例的最大似然估计。它通过使用基因型可能性来解释序列数据中存在的基因型的不确定性,并且不需要来自外部参考小组的个体基因型数据。仿真研究和在实际数据集上的应用表明,我们的方法比以前的方法快得多,并且具有相当的精度。使用来自1000基因组计划的数据,我们表明混合个体的全基因组平均祖先估计在外显子组序列数据和全基因组低覆盖序列数据之间是一致的。最后,我们证明了我们的方法可以用来估计混合比例,使用合并的序列数据,使其成为一个有价值的工具,用于控制群体分层的基于测序的关联研究,利用DNA池。我们的方法是从DNA序列数据估计祖先的有效和通用的工具,可从https://sites.google.com/site/vibansal/software/iAdmix。本文的在线版本(doi:10.1186/s12859-014-0418-7)包含补充材料,可供授权用户使用。
Estimation of individual ancestry from genetic data is useful for the analysis of disease association studies, understanding human population history and interpreting personal genomic variation. New, computationally efficient methods are needed for ancestry inference that can effectively utilize existing information about allele frequencies associated with different human populations and can work directly with DNA sequence reads. We describe a fast method for estimating the relative contribution of known reference populations to an individual’s genetic ancestry. Our method utilizes allele frequencies from the reference populations and individual genotype or sequence data to obtain a maximum likelihood estimate of the global admixture proportions using the BFGS optimization algorithm. It accounts for the uncertainty in genotypes present in sequence data by using genotype likelihoods and does not require individual genotype data from external reference panels. Simulation studies and application of the method to real datasets demonstrate that our method is significantly times faster than previous methods and has comparable accuracy. Using data from the 1000 Genomes project, we show that estimates of the genome-wide average ancestry for admixed individuals are consistent between exome sequence data and whole-genome low-coverage sequence data. Finally, we demonstrate that our method can be used to estimate admixture proportions using pooled sequence data making it a valuable tool for controlling for population stratification in sequencing based association studies that utilize DNA pooling. Our method is an efficient and versatile tool for estimating ancestry from DNA sequence data and is available from https://sites.google.com/site/vibansal/software/iAdmix. The online version of this article (doi:10.1186/s12859-014-0418-7) contains supplementary material, which is available to authorized users.
来自1,092个人基因组的遗传变异的综合图。
DOI: 10.1038/nature11632
发表时间: 2012-11-01
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.3389/fgene.2012.00322
发表时间: 2012
影响因子: 3.7
作者:
Libiger O;Schork NJ
通讯作者: Schork NJ
DOI: 10.1371/journal.pone.0018353
发表时间: 2011-03-30
期刊: PLOS ONE
影响因子: 3.7
作者:
Bansal, Vikas;Tewhey, Ryan;Schork, Nicholas J.
通讯作者: Schork, Nicholas J.
DOI: 10.1186/1471-2105-12-246
发表时间: 2011-06-18
期刊: BMC bioinformatics
影响因子: 3
作者:
Alexander DH;Lange K
通讯作者: Lange K
DOI: 10.1101/gr.078212.108
发表时间: 2008-11-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Li, Heng;Ruan, Jue;Durbin, Richard
通讯作者: Durbin, Richard