Inference of multiple-wave admixtures by length distribution of ancestral tracks

Inference of multiple-wave admixtures by length distribution of ancestral tracks
复制标题

DOI:
10.1038/s41437-017-0041-2
复制
发表时间:
2018-07-01
期刊:
影响因子:
3.8
通讯作者:
Xu, Shuhua
Xu, Shuhua
中科院分区:
生物学2区
文献类型:
--
作者:
Ni, Xumin;Yuan, Kai;Xu, Shuhua

文献摘要

被引文献

相似文献

混合基因组的祖先轨迹对群体历史推断具有重要价值。目前已有一些基于祖先轨迹推断混合历史的方法,但这些方法都存在同样的缺陷,即只能推断某些特定模型下的种群混合历史。此外,如果具体的模型偏离了实际情况,历史的推断可能会有偏差甚至不可靠。为了解决这一问题,我们首先提出了一个通用的离散外加剂模型来描述具有多祖先种群和多波外加剂的外加剂历史。然后推导了一般离散外加剂模型下祖先轨迹的长度分布。我们进一步开发了一种新的方法,MultiWaver,来探索多波混合历史。该方法可以根据祖先轨迹的长度分布自动确定最优外加剂模型,并在此最优模型下估计相应的参数。具体来说,我们使用了似然比检验(LRT)来确定混合波的数量,并实现了期望最大化(EM)算法来估计参数。我们使用仿真研究来验证我们方法的可靠性和有效性。最后,将我们的方法应用于非裔美国人和墨西哥人的真实数据集,取得了良好的效果,并对维吾尔族和哈扎拉族的混合历史有了新的认识。
The ancestral tracks in admixed genomes are valuable for population history inference. While a few methods have been developed to infer admixture history based on ancestral tracks, these methods suffer the same flaw: only population admixture history under some specific models can be inferred. In addition, the inference of history might be biased or even unreliable if the specific model deviates from the real situation. To address this problem, we firstly proposed a general discrete admixture model to describe the admixture history with multiple ancestral populations and multiple-wave admixtures. We next deduced the length distribution of ancestral tracks under the general discrete admixture model. We further developed a new method, MultiWaver, to explore multiple-wave admixture histories. Our method could automatically determine an optimal admixture model based on the length distribution of ancestral tracks, and estimate the corresponding parameters under this optimal model. Specifically, we used a likelihood ratio test (LRT) to determine the number of admixture waves, and implemented an expectation-maximization (EM) algorithm to estimate parameters. We used simulation studies to validate the reliability and effectiveness of our method. Finally, good performance was observed when our method was applied to real data sets of African Americans and Mexicans, and new insights were gained into the admixture history of Uyghurs and Hazaras.