How Robust Are "Isolation with Migration" Analyses to Violations of the IM Model? A Simulation Study

How Robust Are "Isolation with Migration" Analyses to Violations of the IM Model? A Simulation Study
复制标题

DOI:
10.1093/molbev/msp233
复制
发表时间:
2010-02-01
影响因子:
10.7
通讯作者:
Rieseberg, Loren H.
Rieseberg, Loren H.
中科院分区:
生物学1区
文献类型:
--
作者:
Strasburg, Jared L.;Rieseberg, Loren H.

文献摘要

被引文献

相似文献

过去十年开发的方法使得以前所未有的准确性和精度估计分子人口统计参数成为可能,例如有效群体规模、分化时间和基因流。然而,他们对物种历史的某些方面和遗传数据的性质做出了简化的假设,并且尚不清楚它们对违反这些假设的稳健程度如何。在这里,我们使用模拟数据集来检查许多违反“迁移隔离”(IM)模型的影响,包括位点内重组、种群结构、来自未采样物种的基因流、位点之间的联系以及发散选择,对使用程序 IMA 进行的人口统计参数估计的影响。我们还研究了除了 IMA 中可用的两个相对简单的模型之外,拥有适合核苷酸替代模型的数据的效果。我们发现,与现实场景中经常遇到的情况相比,IMA 估计通常对于小到中度违反 IM 模型假设的情况相当稳健。特别是,物种内的种群结构(几乎所有物种都在某种程度上遇到的情况)即使对于相当高水平的结构,对参数估计也几乎没有影响。同样,当数据集被削减为明显非重组的块时,大多数参数估计对于显着的重组水平是稳健的,尽管当包括具有重组的整个数据集时,一些估计会引入显着的偏差。相反,与核苷酸替代模型的拟合不良可能会导致错误率增加,在某些情况下是由于可预测的偏差,而在其他情况下是由于在相同条件下模拟的数据集之间参数估计的方差增加。
Methods developed over the past decade have made it possible to estimate molecular demographic parameters such as effective population size, divergence time, and gene flow with unprecedented accuracy and precision. However, they make simplifying assumptions about certain aspects of the species' histories and the nature of the genetic data, and it is not clear how robust they are to violations of these assumptions. Here, we use simulated data sets to examine the effects of a number of violations of the "Isolation with Migration" (IM) model, including intralocus recombination, population structure, gene flow from an unsampled species, linkage among loci, and divergent selection, on demographic parameter estimates made using the program IMA. We also examine the effect of having data that fit a nucleotide substitution model other than the two relatively simple models available in IMA. We find that IMA estimates are generally quite robust to small to moderate violations of the IM model assumptions, comparable with what is often encountered in real-world scenarios. In particular, population structure within species, a condition encountered to some degree in virtually all species, has little effect on parameter estimates even for fairly high levels of structure. Likewise, most parameter estimates are robust to significant levels of recombination when data sets are pared down to apparently nonrecombining blocks, although substantial bias is introduced to several estimates when the entire data set with recombination is included. In contrast, a poor fit to the nucleotide substitution model can result in an increased error rate, in some cases due to a predictable bias and in other cases due to an increase in variance in parameter estimates among data sets simulated under the same conditions.