Long-branch attraction bias and inconsistency in Bayesian phylogenetics.

Long-branch attraction bias and inconsistency in Bayesian phylogenetics.
复制标题

DOI:
10.1371/journal.pone.0007891
复制
发表时间:
2009-12-09
期刊:
影响因子:
3.7
通讯作者:
Thornton JW
Thornton JW
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Kolaczkowski B;Thornton JW

文献摘要

参考文献

被引文献

相似文献

系统发育关系的贝叶斯推断(BI)使用与其前体最大似然(ML)相同的进化概率模型,因此BI通常被假设为共享ML的理想统计特性,例如在给定准确模型的情况下,拓扑结构的基本无偏推断以及随着数据量的增加而越来越可靠的推断。在这里,我们表明,BI,不像ML,是偏向于有利于拓扑结构,组长分支在一起,即使在一组的进化参数的真实模型和先验分布是已知的。使用实验模拟研究和数值和数学分析,我们表明,这种偏见变得更加严重,因为更多的数据进行分析,导致BI推断一个不正确的树作为最大的后验概率与渐近高的支持序列长度接近无穷大。BI的长分支吸引力的偏见是相对较弱时,真正的模型是简单的,但变得明显时,序列位点的演变不均匀,即使这种复杂性被纳入模型。这种偏见,这是明显的控制下的模拟条件下,并在分析的经验序列数据,也使得BI的效率较低,鲁棒性较低的使用一个不正确的进化模型比ML。令人惊讶的是,BI的偏差是由该方法的一个声明引起的-它通过对可能值的分布进行积分而不是像ML那样从数据中估计它们,从而合并了关于分支长度的不确定性。我们的研究结果表明,使用BI推断的树应谨慎解释,ML可能是现代系统发育分析的一个更可靠的框架。
Bayesian inference (BI) of phylogenetic relationships uses the same probabilistic models of evolution as its precursor maximum likelihood (ML), so BI has generally been assumed to share ML's desirable statistical properties, such as largely unbiased inference of topology given an accurate model and increasingly reliable inferences as the amount of data increases. Here we show that BI, unlike ML, is biased in favor of topologies that group long branches together, even when the true model and prior distributions of evolutionary parameters over a group of phylogenies are known. Using experimental simulation studies and numerical and mathematical analyses, we show that this bias becomes more severe as more data are analyzed, causing BI to infer an incorrect tree as the maximum a posteriori phylogeny with asymptotically high support as sequence length approaches infinity. BI's long branch attraction bias is relatively weak when the true model is simple but becomes pronounced when sequence sites evolve heterogeneously, even when this complexity is incorporated in the model. This bias—which is apparent under both controlled simulation conditions and in analyses of empirical sequence data—also makes BI less efficient and less robust to the use of an incorrect evolutionary model than ML. Surprisingly, BI's bias is caused by one of the method's stated advantages—that it incorporates uncertainty about branch lengths by integrating over a distribution of possible values instead of estimating them from the data, as ML does. Our findings suggest that trees inferred using BI should be interpreted with caution and that ML may be a more reliable framework for modern phylogenetic analysis.
DOI: 10.1080/10635150490522629
发表时间: 2004-12-01
期刊: SYSTEMATIC BIOLOGY
影响因子: 6.5
作者:
Huelsenbeck, JP;Rannala, B
通讯作者: Rannala, B
DOI: 10.1016/j.ympev.2004.06.015
发表时间: 2004-11-01
影响因子: 4.1
作者:
Anderson, FE;Swofford, DL
通讯作者: Swofford, DL
DOI: 10.1109/tac.1974.1100705
发表时间: 1974-01-01
影响因子: 6.8
作者:
AKAIKE, H
通讯作者: AKAIKE, H
DOI: 10.1126/science.1065156
发表时间: 2001-12-14
期刊: SCIENCE
影响因子: 56.9
作者:
Karol, KG;McCourt, RM;Delwiche, CF
通讯作者: Delwiche, CF
DOI: 10.1080/10635150600755453
发表时间: 2006-08-01
期刊: SYSTEMATIC BIOLOGY
影响因子: 6.5
作者:
Anisimova, Maria;Gascuel, Olivier
通讯作者: Gascuel, Olivier