Increasing the efficiency of searches for the maximum likelihood tree in a phylogenetic analysis of up to 150 nucleotide sequences

Increasing the efficiency of searches for the maximum likelihood tree in a phylogenetic analysis of up to 150 nucleotide sequences
复制标题

DOI:
10.1080/10635150701779808
复制
发表时间:
2007-12-01
期刊:
影响因子:
6.5
通讯作者:
Morrison, David A.
Morrison, David A.
中科院分区:
生物学1区
文献类型:
--
作者:
Morrison, David A.

文献摘要

被引文献

相似文献

即使当最大似然(NIL)树是一个更好的估计,真正的系统发育树比其他方法产生的,一个穷人的NIL搜索的结果可能不会比更彻底的搜索在一些更快的标准。因此,找到全局最优NIL树的能力是重要的。在这里,我比较了一系列的启发式搜索策略(及其相关的计算机程序)在定位NIL树的20个经验数据集与14至158个序列和411至120,762对齐核苷酸的成功。三个不同的主题进行了讨论:成功的搜索策略的某些功能的数据,生成的搜索开始树,和探索多个岛屿的树木。作为起始树,基于绝对差异的邻居连接树(包括BioNJ树)、逐步添加的简约树(有或没有最近邻居交换(NNI)分支交换)和逐步添加的NIL树之间的差异很小。后者产生了最好的NIL平均得分,但数量级比替代品慢。BioNJ树平均排名第二。作为搜索策略,星星分解和四元组困惑是最慢的,产生最差的NIL得分。具有默认选项的DPRml、IQPNNI、MultiPhyl、PhyML、PhyNav和Treeplant程序产生了定性相似的结果,当数据集具有低系统发育信息时,每个程序都定位倾向于处于NNI次优(而不是全局最优)的单个树。对于这样的数据集,有多个树岛具有非常相似的NIL分数。对于50个序列包含大约500个比对核苷酸和100个序列包含3,000个核苷酸的数据集,似然表面仅变得相对简单。RAxML和GARLI程序允许轻松探索多个岛屿,但这两个程序也倾向于发现NNI次优。使用PAUP* 的可能性棘轮的新开发版本成功地找到了多个岛屿的峰值,但其速度需要提高。
Even when the maximum likelihood (NIL) tree is a better estimate of the true phylogenetic tree than those produced by other methods, the result of a poor NIL search may be no better than that of a more thorough search under some faster criterion. The ability to find the globally optimal NIL tree is therefore important. Here, I compare a range of heuristic search strategies (and their associated computer programs) in terms of their success at locating the NIL tree for 20 empirical data sets with 14 to 158 sequences and 411 to 120,762 aligned nucleotides. Three distinct topics are discussed: the success of the search strategies in relation to certain features of the data, the generation of starting trees for the search, and the exploration of multiple islands of trees. As a starting tree, there was little difference among the neighbor-joining tree based on absolute differences (including the BioNJ tree), the stepwise-addition parsimony tree (with or without nearest-neighbor-interchange (NNI) branch swapping), and the stepwise-addition NIL tree. The latter produced the best NIL score on average but was orders of magnitude slower than the alternatives. The BioNJ tree was second best on average. As search strategies, star decomposition and quartet puzzling were the slowest and produced the worst NIL scores. The DPRml, IQPNNI, MultiPhyl, PhyML, PhyNav, and TreeFinder programs with default options produced qualitatively similar results, each locating a single tree that tended to be in an NNI suboptimum (rather than the global optimum) when the data set had low phylogenetic information. For such data sets, there were multiple tree islands with very similar NIL scores. The likelihood surface only became relatively simple for data sets that contained approximately 500 aligned nucleotides for 50 sequences and 3,000 nucleotides for 100 sequences. The RAxML and GARLI programs allowed multiple islands to be explored easily, but both programs also tended to find NNI suboptima. A newly developed version of the likelihood ratchet using PAUP* successfully found the peaks of multiple islands, but its speed needs to be improved.