Theoretical foundation of the balanced minimum evolution method of phylogenetic inference and its relationship to weighted least-squares tree fitting

Theoretical foundation of the balanced minimum evolution method of phylogenetic inference and its relationship to weighted least-squares tree fitting
复制标题

DOI:
10.1093/molbev.msh049
复制
发表时间:
2004-03-01
影响因子:
10.7
通讯作者:
Gascuel, O
Gascuel, O
中科院分区:
生物学1区
文献类型:
--
作者:
Desper, R;Gascuel, O

文献摘要

被引文献

相似文献

由于其速度,距离法仍然是在非常大的分类群集合上构建系统发育树的最大希望。最近(R.德斯珀和O.加斯凯尔,《计算生物学杂志》9:687 - 705,2002年),我们引入了一种新的“平衡”最小进化(BME)原理,它基于Y.波普林(《分子进化杂志》51:41 - 47,2000年)的一个分支长度估计方案。初步模拟表明,我们实现BME原理的程序FASTME比我们测试的所有其他距离方法更准确或与之相当,且运行时间明显快于邻接法(NJ)。本文进一步探究了BME原理的特性,并解释和说明了其令人印象深刻的拓扑准确性。我们证明BME原理是加权最小二乘法的一种特殊情况,其距离估计具有生物学意义上的方差。我们表明BME原理在统计上是一致的。我们证明FASTME只产生具有正分支长度的树,这一特征将该方法与NJ(及相关方法)区分开来,因为NJ可能会产生具有生物学上无意义的负长度分支的树。最后,我们考虑一个大型模拟数据集,其中有5000个100个分类群的树,由阿尔多斯β分裂分布生成,涵盖了从尤尔 - 哈丁分布到均匀分布的一系列分布,并使用了一种类似共变密码子的序列进化模型。FASTME生成树的速度比NJ快,且比WEIGHBOR和PAUP*的加权最小二乘法实现快得多。此外,在从尤尔 - 哈丁分布到均匀分布的所有设置以及最大成对差异和偏离分子钟的所有范围中,FASTME树始终更准确。有趣的是,共变密码子参数对任何算法的树质量影响都很小。FASTME可在网上免费获取。
Due to its speed, the distance approach remains the best hope for building phylogenies on very large sets of taxa. Recently (R. Desper and O. Gascuel, J. Comp. Biol. 9:687-705, 2002), we introduced a new "balanced" minimum evolution (BME) principle, based on a branch length estimation scheme of Y. Pauplin (J. Mol. Evol. 51:41-47, 2000). Initial simulations suggested that FASTME, our program implementing the BME principle, was more accurate than or equivalent to all other distance methods we tested, with running time significantly faster than Neighbor-Joining (NJ). This article further explores the properties of the BME principle, and it explains and illustrates its impressive topological accuracy. We prove that the BME principle is a special case of the weighted least-squares approach, with biologically meaningful variances of the distance estimates. We show that the BME principle is statistically consistent. We demonstrate that FASTME only produces trees with positive branch lengths, a feature that separates this approach from NJ (and related methods) that may produce trees with branches with biologically meaningless negative lengths. Finally, we consider a large simulated data set, with 5,000 100-taxon trees generated by the Aldoas beta-splitting distribution encompassing a range of distributions from Yule-Harding to uniform, and using a covarion-like model of sequence evolution. FASTME produces trees faster than NJ, and much faster than WEIGHBOR and the weighted least-squares implementation of PAUP*. Moreover, FASTME trees are consistently more accurate at all settings, ranging from Yule-Harding to uniform distributions, and all ranges of maximum pairwise divergence and departure from molecular clock. Interestingly, the covarion parameter has little effect on the tree quality for any of the algorithms. FASTME is freely available on the web.