Efficiencies of fast algorithms of phylogenetic inference under the criteria of maximum parsimony, minimum evolution, and maximum likelihood when a large number of sequences are used.

Efficiencies of fast algorithms of phylogenetic inference under the criteria of maximum parsimony, minimum evolution, and maximum likelihood when a large number of sequences are used.
复制标题

DOI:
10.1093/oxfordjournals.molbev.a026408
复制
发表时间:
2000-08
影响因子:
10.7
通讯作者:
Kei-ichiro Takahashi;M. Nei
Kei-ichiro Takahashi;M. Nei
中科院分区:
生物学1区
文献类型:
--
作者:
Kei-ichiro Takahashi;M. Nei

文献摘要

被引文献

相似文献

在通过最大简约(MP)、最小进化(ME)和最大似然(ML)方法进行系统发育推断时,通常会对MP、ME和ML树进行广泛的启发式搜索,检查大量不同的拓扑结构。然而,这些广泛的搜索往往会给出不正确的树拓扑。这里我们通过大量的计算机模拟显示,当核苷酸序列的数量(m)是大型和核苷酸(n)使用的数量相对较小,简单的议员或ML树搜索算法等逐步添加(SA) +最近的邻居交换(NNI)搜索和SA +子树修剪再接枝于(SPR)搜索一样高效的搜索算法,如SA +树bisection-reconnection(创业)搜索推断真实的树。在ME方法中,简单邻居连接(NJ)算法与广泛的NJ+TBR搜索一样有效或更有效。我们表明,当使用ME方法时,简单的p距离通常比更复杂的距离测量(如Hasegawa-Kishino-Yano (HKY)距离)在系统发育推断中给出更好的结果,即使核苷酸取代遵循HKY模型。当使用ML方法时,系统发育推断的简单Jukes-Cantor (JC)模型通常比HKY模型表现出更好的性能,即使HKY模型的似然值远高于JC模型。这表明,至少在本案例中,使用似然比检验或AIC指数选择替代模型是不合适的。当n相对于m较小且序列发散程度较高时,具有p距离的NJ方法往往比具有JC模型的ML方法表现出更好的性能。然而,当序列散度水平较低时,情况并非如此。
In phylogenetic inference by maximum-parsimony (MP), minimum-evolution (ME), and maximum-likelihood (ML) methods, it is customary to conduct extensive heuristic searches of MP, ME, and ML trees, examining a large number of different topologies. However, these extensive searches tend to give incorrect tree topologies. Here we show by extensive computer simulation that when the number of nucleotide sequences (m) is large and the number of nucleotides used (n) is relatively small, the simple MP or ML tree search algorithms such as the stepwise addition (SA) plus nearest neighbor interchange (NNI) search and the SA plus subtree pruning regrafting (SPR) search are as efficient as the extensive search algorithms such as the SA plus tree bisection-reconnection (TBR) search in inferring the true tree. In the case of ME methods, the simple neighbor-joining (NJ) algorithm is as efficient as or more efficient than the extensive NJ+TBR search. We show that when ME methods are used, the simple p distance generally gives better results in phylogenetic inference than more complicated distance measures such as the Hasegawa-Kishino-Yano (HKY) distance, even when nucleotide substitution follows the HKY model. When ML methods are used, the simple Jukes-Cantor (JC) model of phylogenetic inference generally shows a better performance than the HKY model even if the likelihood value for the HKY model is much higher than that for the JC model. This indicates that at least in the present case, selecting of a substitution model by using the likelihood ratio test or the AIC index is not appropriate. When n is small relative to m and the extent of sequence divergence is high, the NJ method with p distance often shows a better performance than ML methods with the JC model. However, when the level of sequence divergence is low, this is not the case.