Bayesian phylogenetic model selection using reversible jump Markov chain Monte Carlo

Bayesian phylogenetic model selection using reversible jump Markov chain Monte Carlo
复制标题

DOI:
10.1093/molbev/msh123
复制
发表时间:
2004-06-01
影响因子:
10.7
通讯作者:
Alfaro, ME
Alfaro, ME
中科院分区:
生物学1区
文献类型:
--
作者:
Huelsenbeck, JP;Larget, B;Alfaro, ME

文献摘要

被引文献

相似文献

在分子遗传学中,一个常见的问题是选择一个DNA替换模型,该模型能够很好地解释DNA序列比对,而不引入多余的参数。已经使用了许多方法来在一小组候选替代模型中进行选择,例如似然比检验、赤池信息准则(AIC)、贝叶斯信息准则(BIC)和贝叶斯因子。这些标准中的任何一个的当前实现都受到以下限制:仅检查一小部分模型,或者测试不允许对非嵌套模型进行容易的比较。在这篇文章中,我们扩展了候选替代模型池,以包括所有可能的时间可逆模型。这一组包括已经描述的七个模型。我们展示了如何贝叶斯因子可以计算这些模型使用可逆跳马尔可夫链蒙特卡罗,并将该方法应用于16个DNA序列比对。对于每个数据集,我们将具有最佳贝叶斯因子的模型与使用AIC和BIC选择的最佳模型进行比较。我们发现,在任何这些标准下的最佳模型不一定是最复杂的模型;具有中间数量的替代类型的模型通常做得最好。此外,几乎所有被选为最佳模型的模型都没有将转换率限制为与颠换率相同,这表明转换/颠换率偏差在决定选择哪些模型方面起着最大的作用。重要的是,这里描述的可逆跳马尔可夫链蒙特卡罗算法允许估计的同源性(和其他系统发育模型参数)进行,同时考虑到DNA取代模型中的不确定性。
A common problem in molecular phylogenetics is choosing a model of DNA substitution that does a good job of explaining the DNA sequence alignment without introducing superfluous parameters. A number of methods have been used to choose among a small set of candidate substitution models, such as the likelihood ratio test, the Akaike Information Criterion (AIC), tire Bayesian Information Criterion (BIC), and Bayes factors. Current implementations of any of these criteria suffer from the limitation that only a small set of models are examined, or that the test does not allow easy comparison of non-nested models. In this article, we expand the pool of candidate substitution models to include all possible time-reversible models. This set includes seven models that have already been described. We show how Bayes factors can be calculated for these models using reversible jump Markov chain Monte Carlo, and apply the method to 16 DNA sequence alignments. For each data set, we compare the model with the best Bayes factor to the best models chosen using AIC and BIC. We find that the best model under any of these criteria is not necessarily the most complicated one; models with an intermediate number of substitution types typically do best. Moreover, almost all of the models that are chosen as best do not constrain a transition rate to be the same as a transversion rate, suggesting that it is the transition/transversion rate bias that plays the largest role in determining which models are selected. Importantly, the reversible jump Markov chain Monte Carlo algorithm described here allows estimation of phylogeny (and other phylogenetic model parameters) to be performed while accounting for uncertainty in the model of DNA substitution.