Improving the Accuracy of Demographic and Molecular Clock Model Comparison While Accommodating Phylogenetic Uncertainty

Improving the Accuracy of Demographic and Molecular Clock Model Comparison While Accommodating Phylogenetic Uncertainty
复制标题

DOI:
10.1093/molbev/mss084
复制
发表时间:
2012-09-01
影响因子:
10.7
通讯作者:
Alekseyenko, Alexander V.
Alekseyenko, Alexander V.
中科院分区:
生物学1区
文献类型:
--
作者:
Baele, Guy;Lemey, Philippe;Alekseyenko, Alexander V.

文献摘要

被引文献

相似文献

贝叶斯遗传学和分子进化领域中用于模型选择的边缘似然估计的最新发展强调了调和均值估计(HME)的性能不佳。虽然这些研究已经显示了适用于标准正态分布的例子和小的真实世界的数据集的新方法的优点,目前还不知道很多关于这些方法的性能和计算问题时,拟合复杂的进化和种群遗传模型的经验真实世界的数据集。此外,这些方法还没有看到在该领域中的广泛应用,由于在常用的系统发育包中缺乏这些计算要求高的技术的实现。我们在这里调查的性能,这些新的边际似然估计,特别是,路径抽样(PS)和垫脚石(SS)抽样比较模型的人口变化和放松分子时钟,使用合成数据和现实世界的例子,其中意外的推断,使用HME。考虑到PS和SS采样的计算需求急剧增加,我们还调查了一个后验模拟为基础的模拟赤池的信息准则(AIC)通过马尔可夫链蒙特卡罗(MCMC),模型比较的方法,共享与HME的吸引力的功能,具有较低的计算开销超过原来的MCMC分析。我们证实,HME系统高估的边际可能性,并未能产生可靠的模型分类,并表明AICM表现更好,可能是一个有用的初始评估模型的选择,但它也是,在较小的程度上,不可靠。我们表明,PS和SS采样大大优于这些估计,并调整有关以前的分析,我们重新分析的三个现实世界的数据集的结论。本文中使用的方法现在可以在BEAST中使用,BEAST是一个强大的用户友好的软件包,用于执行贝叶斯进化分析。
Recent developments in marginal likelihood estimation for model selection in the field of Bayesian phylogenetics and molecular evolution have emphasized the poor performance of the harmonic mean estimator (HME). Although these studies have shown the merits of new approaches applied to standard normally distributed examples and small real-world data sets, not much is currently known concerning the performance and computational issues of these methods when fitting complex evolutionary and population genetic models to empirical real-world data sets. Further, these approaches have not yet seen widespread application in the field due to the lack of implementations of these computationally demanding techniques in commonly used phylogenetic packages. We here investigate the performance of some of these new marginal likelihood estimators, specifically, path sampling (PS) and stepping-stone (SS) sampling for comparing models of demographic change and relaxed molecular clocks, using synthetic data and real-world examples for which unexpected inferences were made using the HME. Given the drastically increased computational demands of PS and SS sampling, we also investigate a posterior simulation-based analogue of Akaike's information criterion (AIC) through Markov chain Monte Carlo (MCMC), a model comparison approach that shares with the HME the appealing feature of having a low computational overhead over the original MCMC analysis. We confirm that the HME systematically overestimates the marginal likelihood and fails to yield reliable model classification and show that the AICM performs better and may be a useful initial evaluation of model choice but that it is also, to a lesser degree, unreliable. We show that PS and SS sampling substantially outperform these estimators and adjust the conclusions made concerning previous analyses for the three real-world data sets that we reanalyzed. The methods used in this article are now available in BEAST, a powerful user-friendly software package to perform Bayesian evolutionary analyses.