Selecting the best-fit model of nucleotide substitution

Selecting the best-fit model of nucleotide substitution
复制标题

DOI:
10.1080/106351501750435121
复制
发表时间:
2001-08-01
期刊:
影响因子:
6.5
通讯作者:
Crandall, KA
Crandall, KA
中科院分区:
生物学1区
文献类型:
--
作者:
Posada, D;Crandall, KA

文献摘要

被引文献

相似文献

尽管核苷酸取代模型在遗传学中的相关作用,但在不同模型中进行选择仍然是一个问题。已经提出了几种统计方法来选择最适合手头数据的模型,但它们的绝对和相对性能尚未得到表征。在这项研究中,我们比较在各种条件下的性能不同的层次和动态似然比检验,赤池和贝叶斯信息的方法,选择最适合的模型的核苷酸取代。我们专门研究的拓扑结构的作用,用于估计不同模型的可能性和假设进行测试的顺序的重要性。我们通过在已知的核苷酸替换模型下模拟DNA序列并记录不同方法恢复真实模型的频率来做到这一点。我们的研究结果表明,模型选择是相当准确的,并表明,一些似然比检验方法的整体表现优于赤池或贝叶斯信息标准。用于估计似然分数的树不影响模型选择,除非它是随机选择的树。在某些情况下,假设检验的顺序以及检验序列中初始模型的复杂性会影响模型的选择。模型拟合在系统发育学中的应用已经提出了很多年,但许多作者仍然任意选择他们的模型,通常使用标准计算机程序中实现的默认模型进行系统发育估计。我们在这里表明,最佳拟合模型可以很容易地确定。因此,考虑到模型的相关性,模型拟合应该是任何使用进化模型的系统发育分析的常规。
Despite the relevant role of models of nucleotide substitution in phylogenetics, choosing among different models remains a problem. Several statistical methods for selecting the model that best fits the data at hand have been proposed, but their absolute and relative performance has not yet been characterized. In this study, we compare under various conditions the performance of different hierarchical and dynamic likelihood ratio tests, and of Akaike and Bayesian information methods, for selecting best-fit models of nucleotide substitution. We specifically examine the role of the topology used to estimate the likelihood of the different models and the importance of the order in which hypotheses are tested. We do this by simulating DNA sequences under a known model of nucleotide substitution and recording how often this true model is recovered by the different methods. Our results suggest that model selection is reasonably accurate and indicate that some likelihood ratio test methods perform overall better than the Akaike or Bayesian information criteria. The tree used to estimate the likelihood scores does not influence model selection unless it is a randomly chosen tree. The order in which hypotheses are tested, and the complexity of the initial model in the sequence of tests, influence model selection in some cases. Model fitting in phylogenetics has been suggested for many years, yet many authors still arbitrarily choose their models, often using the default models implemented in standard computer programs for phylogenetic estimation. We show here that a best-fit model can be readily identified. Consequently, given the relevance of models, model fitting should be routine in any phylogenetic analysis that uses models of evolution.