Prospects for inferring very large phylogenies by using the neighbor-joining method

Prospects for inferring very large phylogenies by using the neighbor-joining method
复制标题

DOI:
10.1073/pnas.0404206101
复制
发表时间:
2004-07-27
影响因子:
11.1
通讯作者:
Kumar, S
Kumar, S
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Tamura, K;Nei, M;Kumar, S

文献摘要

被引文献

相似文献

目前重建多基因家族的生命树和历史的努力需要推断由数千个基因序列组成的系统发育。然而,对于这样大的数据集,即使是适度的探索,需要确定最佳的树的树空间几乎是不可能的。对于这些情况下,相邻连接(NJ)的方法是经常使用的,因为它证明了较小的数据集的准确性和计算速度。然而,随着数据集的增长,NJ算法所检查的树空间的比例变得微不足道。在这里,我们报告的结果,我们的计算机模拟检查的准确性,NJ树推断非常大的heterogenies。首先,我们提出了一个可能性的方法,同时估计所有成对的距离,通过使用生物现实模型的核苷酸取代。使用这种方法可以纠正高达60%的NJ树错误。我们的模拟结果表明,当使用的序列数从32增加到4,096(128倍)时,即使在谱系之间进化速率存在广泛变化或核苷酸组成和转换/颠换比存在显著偏差的情况下,NJ树的准确性也仅下降约5%。我们的研究结果鼓励使用复杂的模型的核苷酸取代估计进化的距离和暗示光明的前景应用的NJ和相关的方法在推断大的同源性。
Current efforts to reconstruct the tree of life and histories of multigene families demand the inference of phylogenies consisting of thousands of gene sequences. However, for such large data sets even a moderate exploration of the tree space needed to identify the optimal tree is virtually impossible. For these cases the neighbor-joining (NJ) method is frequently used because of its demonstrated accuracy for smaller data sets and its computational speed. As data sets grow, however, the fraction of the tree space examined by the NJ algorithm becomes minuscule. Here, we report the results of our computer simulation for examining the accuracy of NJ trees for inferring very large phylogenies. First we present a likelihood method for the simultaneous estimation of all pairwise distances by using biologically realistic models of nucleotide substitution. Use of this method corrects up to 60% of NJ tree errors. Our simulation results show that the accuracy of NJ trees decline only by approximate to5% when the number of sequences used increases from 32 to 4,096 (128 times) even in the presence of extensive variation in the evolutionary rate among lineages or significant biases in the nucleotide composition and transition/transversion ratio. Our results encourage the use of complex models of nucleotide substitution for estimating evolutionary distances and hint at bright prospects for the application of the NJ and related methods in inferring large phylogenies.