MultiWaver 2.0: modeling discrete and continuous gene flow to reconstruct complex population admixtures

MultiWaver 2.0: modeling discrete and continuous gene flow to reconstruct complex population admixtures
复制标题

DOI:
10.1038/s41431-018-0259-3
复制
发表时间:
2019-01-01
影响因子:
5.2
通讯作者:
Xu, Shuhua
Xu, Shuhua
中科院分区:
生物学2区
文献类型:
--
作者:
Ni, Xumin;Yuan, Kai;Xu, Shuhua

文献摘要

被引文献

相似文献

我们开发MultiWaver软件系列的目标是能够在各种复杂场景下推断种群混合历史。早期版本的MultiWaver只考虑离散混合模型。在这里,我们报告了一个新开发的版本,MultiWaver 2.0,实现了一个更灵活的框架,并能够推断多波混合历史下的离散和连续混合模型。MultiWaver 2.0可以根据染色体祖先径迹的长度分布自动选择最佳混合模型,并在所选模型下估计相应的参数。具体而言,对于离散混合模型,我们使用似然比检验(LRT)来确定最佳离散模型和期望最大化算法来估计参数。此外,根据贝叶斯信息准则(BIC)的原理,我们将最优离散模型与几种连续混合模型进行了比较。在MultiWaver 2.0中,我们还应用了自举技术为所选模型提供支持水平和混合时间估计的置信区间(CI)。仿真研究验证了该方法的可靠性和有效性。最后,该程序在应用于典型混合人群的真实的数据集时表现良好,如非洲裔美国人,维吾尔族和哈扎拉人。
Our goal in developing the MultiWaver software series was to be able to infer population admixture history under various complex scenarios. The earlier version of MultiWaver considered only discrete admixture models. Here, we report a newly developed version, MultiWaver 2.0, that implements a more flexible framework and is capable of inferring multiple-wave admixture histories under both discrete and continuous admixture models. MultiWaver 2.0 can automatically select an optimal admixture model based on the length distribution of ancestral tracks of chromosomes, and the program can estimate the corresponding parameters under the selected model. Specifically, for discrete admixture models, we used a likelihood ratio test (LRT) to determine the optimal discrete model and an expectation-maximization algorithm to estimate the parameters. In addition, according to the principles of the Bayesian Information Criterion (BIC), we compared the optimal discrete model with several continuous admixture models. In MultiWaver 2.0, we also applied a bootstrapping technique to provide levels of support for the chosen model and the confidence interval (CI) of the estimations of admixture time. Simulation studies validated the reliability and effectiveness of our method. Finally, the program performed well when applied to real datasets of typical admixed populations, such as African Americans, Uyghurs, and Hazaras.