Model misspecification and probabilistic tests of topology: Evidence from empirical data sets

Model misspecification and probabilistic tests of topology: Evidence from empirical data sets
复制标题

DOI:
10.1080/10635150290069922
复制
发表时间:
2002-05-01
期刊:
影响因子:
6.5
通讯作者:
Buckley, TR
Buckley, TR
中科院分区:
生物学1区
文献类型:
--
作者:
Buckley, TR

文献摘要

被引文献

相似文献

拓扑结构的概率检验为评估相互竞争的系统发育假说提供了一种强有力的手段。对五个数据集探讨了非参数的Shimodaira - Hasegawa(SH)检验、参数的Swofford - Olsen - Waddell - Hillis(SOWH)检验以及贝叶斯后验概率的表现,对于这五个数据集,所有的系统发育关系都具有非常高的确定性。这些结果与先前的模拟研究一致,先前的研究表明,由于模型误设以及分支长度异质性,SOWH检验容易产生第一类错误。这些结果还表明,当零假设实际上正确时,SOWH检验可能对真实拓扑结构过度自信。相比之下,观察到SH检验要保守得多,即使在高替代率和分支长度异质性的情况下也是如此。对于那些SOWH检验具有误导性的一些数据集,贝叶斯后验概率也具有误导性。所有检验的结果都受到确切的替代模型假设的强烈影响。简单的模型,尤其是那些假设位点间速率同质的模型,具有更高的第一类错误率,并且更有可能产生误导性的后验概率。对于其中一些数据集,常用的替代模型似乎不足以用SOWH检验和贝叶斯方法估计适当的不确定性水平。讨论了两种最大似然检验之间统计功效差异的原因,并与贝叶斯方法进行了对比。
Probabilistic tests of topology offer a powerful means of evaluating competing phylogenetic hypotheses. The performance of the nonparametric Shimodaira-Hasegawa (SH) test, the parametric Swofford-Olsen-Waddell-Hillis (SOWH) test, and Bayesian posterior probabilities were explored for five data sets for which all the phylogenetic relationships are known with a very high degree of certainty. These results are consistent with previous simulation studies that have indicated a tendency for the SOWH test to be prone to generating Type 1 errors because of model misspecification coupled with branch length heterogeneity. These results also suggest that the SOWH test may accord overconfidence in the true topology when the null hypothesis is in fact correct. In contrast, the SH test was observed to be much more conservative, even under high substitution rates and branch length heterogeneity. For some of those data sets where the SOWH test proved misleading, the Bayesian posterior probabilities were also misleading. The results of all tests were strongly influenced by the exact substitution model assumptions. Simple models, especially those that assume rate homogeneity among sites, had a higher Type 1 error rate and were more likely to generate misleading posterior probabilities. For some of these data sets, the commonly used substitution models appear to be inadequate for estimating appropriate levels of uncertainty with the SOWH test and Bayesian methods. Reasons for the differences in statistical power between the two maximum likelihood tests are discussed and are contrasted with the Bayesian approach.