Bayesian inference of population size history from multiple loci

Bayesian inference of population size history from multiple loci
复制标题

DOI:
10.1186/1471-2148-8-289
复制
发表时间:
2008-10-23
影响因子:
3.4
通讯作者:
Drummond, Alexei J.
Drummond, Alexei J.
中科院分区:
生物学2区
文献类型:
--
作者:
Heled, Joseph;Drummond, Alexei J.

文献摘要

被引文献

相似文献

背景:有效群体大小(Ne)与遗传变异有关,是许多群体遗传学模型的基本参数。自1982年JFC Kingman引入n-coalescent以来,已经开发了许多从遗传数据推断当前和过去种群规模的方法。在这里,我们提出了扩展贝叶斯天际线图,这是一种非参数贝叶斯马尔可夫链蒙特卡罗算法,它在几个方面扩展了以前基于聚结的方法,包括分析多个位点的能力。结果:通过广泛的模拟,我们展示了推断种群规模作为数据量函数的准确性和局限性,包括恢复有关进化瓶颈的信息。我们还分析了两个真实数据集来证明新方法的行为;从埃及取样的单基因丙型肝炎病毒数据集和代表16个不同种群的10个位点的果蝇数据集。结论:多基因座对恢复种群大小动态具有重要作用。来自少数个体的多位点数据可以精确地恢复过去的种群规模瓶颈,这是单位点分析无法表征的。我们还证明了序列数据质量是重要的,因为即使是中等水平的测序错误也会导致对实际水平的群体遗传变异的估计精度大大降低。
Background: Effective population size (Ne) is related to genetic variability and is a basic parameter in many models of population genetics. A number of methods for inferring current and past population sizes from genetic data have been developed since JFC Kingman introduced the n-coalescent in 1982. Here we present the Extended Bayesian Skyline Plot, a non-parametric Bayesian Markov chain Monte Carlo algorithm that extends a previous coalescent-based method in several ways, including the ability to analyze multiple loci.Results: Through extensive simulations we show the accuracy and limitations of inferring population size as a function of the amount of data, including recovering information about evolutionary bottlenecks. We also analyzed two real data sets to demonstrate the behavior of the new method; a single gene Hepatitis C virus data set sampled from Egypt and a 10 locus Drosophila ananassae data set representing 16 different populations.Conclusion: The results demonstrate the essential role of multiple loci in recovering population size dynamics. Multi-locus data from a small number of individuals can precisely recover past bottlenecks in population size which can not be characterized by analysis of a single locus. We also demonstrate that sequence data quality is important because even moderate levels of sequencing errors result in a considerable decrease in estimation accuracy for realistic levels of population genetic variability.