Comparative performance of Bayesian and AIC-based measures of phylogenetic model uncertainty

Comparative performance of Bayesian and AIC-based measures of phylogenetic model uncertainty
复制标题

DOI:
10.1080/10635150500433565
复制
发表时间:
2006-02-01
期刊:
影响因子:
6.5
通讯作者:
Huelsenbeck, JP
Huelsenbeck, JP
中科院分区:
生物学1区
文献类型:
--
作者:
Alfaro, ME;Huelsenbeck, JP

文献摘要

被引文献

相似文献

可逆跳跃马尔可夫链蒙特卡罗(RJ-MCMC)是一种同时评估多个相关(但不一定嵌套)统计模型的技术,最近被应用于系统发育模型选择问题。在这里,我们使用模拟方法来评估该方法的性能,并将其与赤池权重进行比较,赤池权重是基于赤池信息准则的模型不确定性的度量。在候选模型假设与生成条件相匹配的条件下,贝叶斯方法和基于aic的方法都表现良好。95%可信区间包含了接近95%的时间生成模型。然而,可信区间的大小不同,贝叶斯可信集包含的模型比基于aic的可信区间少约25%至50%。当所有假设都满足时,后验概率比赤池权重更好地反映了正确的模型,但当某些模型假设被违反时,两种方法的表现相似。贝叶斯后验分布的模型在参数数量上与生成模型更相似,在复杂性上偏差更小。相比之下,akaike加权模型与生成模型的距离更远,并且偏向于略高的复杂性。基于aic的可信区间对速率均匀性假设的违反表现出更强的鲁棒性。AIC和贝叶斯方法都表明,系统发育分析模型的选择可能伴随着大量的不确定性,这表明在分析系统发育数据时应该检查备选模型。
Reversible-jump Markov chain Monte Carlo (RJ-MCMC) is a technique for simultaneously evaluating multiple related (but not necessarily nested) statistical models that has recently been applied to the problem of phylogenetic model selection. Here we use a simulation approach to assess the performance of this method and compare it to Akaike weights, a measure of model uncertainty that is based on the Akaike information criterion. Under conditions where the assumptions of the candidate models matched the generating conditions, both Bayesian and AIC-based methods perform well. The 95% credible interval contained the generating model close to 95% of the time. However, the size of the credible interval differed with the Bayesian credible set containing approximately 25% to 50% fewer models than an AIC-based credible interval. The posterior probability was a better indicator of the correct model than the Akaike weight when all assumptions were met but both measures performed similarly when some model assumptions were violated. Models in the Bayesian posterior distribution were also more similar to the generating model in their number of parameters and were less biased in their complexity. In contrast, Akaike-weighted models were more distant from the generating model and biased towards slightly greater complexity. The AIC-based credible interval appeared to be more robust to the violation of the rate homogeneity assumption. Both AIC and Bayesian approaches suggest that substantial uncertainty can accompany the choice of model for phylogenetic analyses, suggesting that alternative candidate models should be examined in analysis of phylogenetic data.