Fast model-based estimation of ancestry in unrelated individuals

Fast model-based estimation of ancestry in unrelated individuals
复制标题

DOI:
10.1101/gr.094052.109
复制
发表时间:
2009-09-01
期刊:
影响因子:
7
通讯作者:
Lange, Kenneth
Lange, Kenneth
中科院分区:
生物学1区
文献类型:
--
作者:
Alexander, David H.;Novembre, John;Lange, Kenneth

文献摘要

被引文献

相似文献

长期以来,种群分层一直被认为是遗传关联研究中的一个混杂因素。从多基因座基因数据中得出的估计祖先可以用来对种群分层进行统计校正。一种流行的祖先估计技术是广泛应用的程序结构所体现的基于模型的方法。另一种方法是在EIGENSTRAT程序中实现的,它依赖于主成分分析而不是基于模型的估计,并且不直接提供混合分数。EIGENSTRAT之所以广受欢迎,部分原因是它的速度比结构快得多。我们提出了一个新的算法和一个程序,混合,以模型为基础的估计无关的个人的祖先。外加剂采用嵌入结构的似然模型。然而,混合料运行得相当快,解决问题只需几分钟,而这需要几个小时的结构时间。在我们的许多实验中,我们发现混合物的速度几乎和本征速度一样快。混合算法的运行时间改进依赖于使用序列二次规划进行块更新的快速块松弛方案,以及一种新的拟牛顿收敛加速。我们的算法也比程序Frappe中包含的期望最大化(EM)算法的实现更快、更准确。我们的模拟表明,混合对潜在混合系数和祖先等位基因频率的最大似然估计与结构的贝叶斯估计一样准确。在真实世界的数据集上,混合体的估计直接与Structure和EIGENSTRAT的估计相当。综上所述,我们的结果表明,混合的计算速度打开了在基于模型的祖先估计中使用更大的标记集的可能性,并且其估计适合用于关联研究中的人口分层校正。
Population stratification has long been recognized as a confounding factor in genetic association studies. Estimated ancestries, derived from multi-locus genotype data, can be used to perform a statistical correction for population stratification. One popular technique for estimation of ancestry is the model-based approach embodied by the widely applied program structure. Another approach, implemented in the program EIGENSTRAT, relies on Principal Component Analysis rather than model-based estimation and does not directly deliver admixture fractions. EIGENSTRAT has gained in popularity in part owing to its remarkable speed in comparison to structure. We present a new algorithm and a program, ADMIXTURE, for model-based estimation of ancestry in unrelated individuals. ADMIXTURE adopts the likelihood model embedded in structure. However, ADMIXTURE runs considerably faster, solving problems in minutes that take structure hours. In many of our experiments, we have found that ADMIXTURE is almost as fast as EIGENSTRAT. The runtime improvements of ADMIXTURE rely on a fast block relaxation scheme using sequential quadratic programming for block updates, coupled with a novel quasi-Newton acceleration of convergence. Our algorithm also runs faster and with greater accuracy than the implementation of an Expectation-Maximization ( EM) algorithm incorporated in the program FRAPPE. Our simulations show that ADMIXTURE's maximum likelihood estimates of the underlying admixture coefficients and ancestral allele frequencies are as accurate as structure's Bayesian estimates. On real-world data sets, ADMIXTURE's estimates are directly comparable to those from structure and EIGENSTRAT. Taken together, our results show that ADMIXTURE's computational speed opens up the possibility of using a much larger set of markers in model-based ancestry estimation and that its estimates are suitable for use in correcting for population stratification in association studies.