RAxML and FastTree: comparing two methods for large-scale maximum likelihood phylogeny estimation.

RAxML and FastTree: comparing two methods for large-scale maximum likelihood phylogeny estimation.
复制标题

DOI:
10.1371/journal.pone.0027731
复制
发表时间:
2011
期刊:
影响因子:
3.7
通讯作者:
Warnow T
Warnow T
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Liu K;Linder CR;Warnow T

文献摘要

参考文献

被引文献

相似文献

概率估计的统计方法,特别是最大似然法(ML),具有很高的精度和优良的理论特性。然而,RAxML是目前用于大规模ML估计的领先方法,当用于具有数千个分子序列的数据集时,可能需要数周或更长时间。更快的ML估计方法,其中包括FastTree,也已经开发出来,但它们与RAxML的相对性能尚未完全了解。在这项研究中,我们探索了FastTree和RAxML在ML评分、运行时间和拓扑准确性方面的性能,这些性能基于数千个比对(基于模拟和生物核苷酸数据集),最多可达27,634个序列。我们发现,当RAxML和FastTree被限制在相同的运行时间,FastTree产生拓扑更准确的树在几乎所有的情况下。我们还发现,当RAxML被允许运行到完成时,它在ML得分方面比FastTree具有优势,但不会产生更准确的树拓扑。有趣的是,使用FastTree和RAxML计算的树的相对准确性部分取决于序列比对的准确性和数据集大小,因此FastTree在比对相对不准确的大型数据集上比RAxML更准确。最后,RAxML和FastTree的运行时间有很大的不同,所以当运行完成时,RAxML可能比FastTree长几个数量级。因此,我们的研究表明,与RAxML相比,使用FastTree可以非常快速地估计非常大的重复性,树的准确性几乎没有(在某些情况下没有)下降。
Statistical methods for phylogeny estimation, especially maximum likelihood (ML), offer high accuracy with excellent theoretical properties. However, RAxML, the current leading method for large-scale ML estimation, can require weeks or longer when used on datasets with thousands of molecular sequences. Faster methods for ML estimation, among them FastTree, have also been developed, but their relative performance to RAxML is not yet fully understood. In this study, we explore the performance with respect to ML score, running time, and topological accuracy, of FastTree and RAxML on thousands of alignments (based on both simulated and biological nucleotide datasets) with up to 27,634 sequences. We find that when RAxML and FastTree are constrained to the same running time, FastTree produces topologically much more accurate trees in almost all cases. We also find that when RAxML is allowed to run to completion, it provides an advantage over FastTree in terms of the ML score, but does not produce substantially more accurate tree topologies. Interestingly, the relative accuracy of trees computed using FastTree and RAxML depends in part on the accuracy of the sequence alignment and dataset size, so that FastTree can be more accurate than RAxML on large datasets with relatively inaccurate alignments. Finally, the running times of RAxML and FastTree are dramatically different, so that when run to completion, RAxML can take several orders of magnitude longer than FastTree to complete. Thus, our study shows that very large phylogenies can be estimated very quickly using FastTree, with little (and in some cases no) degradation in tree accuracy, as compared to RAxML.
DOI: 10.1093/molbev/msp077
发表时间: 2009-07
影响因子: 10.7
作者:
Price MN;Dehal PS;Arkin AP
通讯作者: Arkin AP
DOI: 10.1093/bioinformatics/14.2.157
发表时间: 1998-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Stoye, J;Evers, D;Meyer, F
通讯作者: Meyer, F
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1093/bioinformatics/btl592
发表时间: 2007-02-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Katoh, Kazutaka;Toh, Hiroyuki
通讯作者: Toh, Hiroyuki
DOI: 10.1126/science.1170540
发表时间: 2009-04-17
期刊: Science (New York, N.Y.)
影响因子: --
作者:
Liu K;Victora GD;Schwickert TA;Guermonprez P;Meredith MM;Yao K;Chu FF;Randolph GJ;Rudensky AY;Nussenzweig M
通讯作者: Nussenzweig M