SUCCESS OF PHYLOGENETIC METHODS IN THE 4-TAXON CASE

SUCCESS OF PHYLOGENETIC METHODS IN THE 4-TAXON CASE
复制标题

DOI:
10.2307/2992463
复制
发表时间:
1993-09-01
期刊:
影响因子:
6.5
通讯作者:
HILLIS, DM
HILLIS, DM
中科院分区:
生物学1区
文献类型:
--
作者:
HUELSENBECK, JP;HILLIS, DM

文献摘要

被引文献

相似文献

用一致性和模拟分析的方法检验了16种系统发育推断方法的成功。成功--造树方法正确识别真正的系统发育的频率--是对一棵无根的四分类树进行检验的。在这项研究中,研究了在大量分枝长度条件下和在三种序列进化模式下的树制作方法。绘制结果图是为了便于各种方法之间的比较。一致性分析表明,对于无限大的样本量,哪些方法收敛到正确的树上。一般简约性、平移简约性和加权简约性在所检查的图空间的部分上是不一致的,尽管不一致的区域是不同的。当序列进化模型与不变量方法的假设相匹配时,莱克的不变量方法一致地估计了所有图形空间的系统发育。然而,当不变量方法的一个假设被违反时,Lake的不变量方法在图空间的很大一部分上变得不一致。一般来说,当距离修正的假设与用于生成模型树的进化模型相匹配时,距离方法(邻接法、加权最小二乘法和未加权最小二乘法)在所检查的所有图形空间上一致地估计系统发育。当距离方法的假设被违反时,这些方法在图空间的部分上变得不一致。无论使用哪种距离,UPGMA在大范围的图形空间上都是不一致的。仿真分析表明,在给定有限数量的字符数据的情况下,树的制作方法是如何执行的。在某些情况下,模拟结果与一致性分析的定量结果不同。一致性分析表明,在某些条件下,Lake的不变量方法在所有图空间上是一致的,而模拟分析表明,Lake的不变量方法在大多数图空间上的表现很差,最多可达500个可变特征。简约法、邻接法和最小二乘法在性状变化量和枝长变化量有限的条件下表现良好。通过对进化较慢的字符进行加权或使用针对多个替换事件进行校正的距离,可以减少造树方法具有误导性的区域。只有对缓慢演变的字符(例如,颠倒简约、加权简约)赋予更高的权重,才能在高变化率下获得良好的性能。只有当分支长度接近时,UPGMA才表现良好。
The success of 16 methods of phylogenetic inference was examined using consistency and simulation analysis. Success-the frequency with which a tree-making method correctly identified the true phylogeny-was examined for an unrooted four-taxon tree. In this study, tree-making methods were examined under a large number of branch-length conditions and under three models of sequence evolution. The results are plotted to facilitate comparisons among the methods. The consistency analysis indicated which methods converge on the correct tree given infinite sample size. General parsimony, transversion parsimony, and weighted parsimony are inconsistent over portions of the graph space examined, although the area of inconsistency varied. Lake's method of invariants consistently estimated phylogeny over all of the graph space when the model of sequence evolution matched the assumptions of the invariants method. However, when one of the assumptions of the invariants method was violated, Lake's method of invariants became inconsistent over a large portion of the graph space. In general, the distance methods (neighbor joining, weighted least squares, and unweighted least squares) consistently estimated phylogeny over all of the graph space examined when the assumptions of the distance correction matched the model of evolution used to generate the model trees. When the assumptions of the distance methods were violated, the methods became inconsistent over portions of the graph space. UPGMA was inconsistent over a large area of the graph space, no matter which distance was used. The simulation analysis showed how tree-making methods perform given limited numbers of character data. In some instances, the simulation results differed quantitatively from the consistency analysis. The consistency analysis indicated that Lake's method of invariants was consistent over all of the graph space under some conditions, whereas the simulation analysis showed that Lake's method of invariants performs poorly over most of the graph space for up to 500 variable characters. Parsimony, neighbor-joining, and the least-squares methods performed well under conditions of limited amount of character change and branch-length variation. By weighting the more slowly evolving characters or using distances that correct for multiple substitution events, the area in which tree-making methods are misleading can be reduced. Good performance at high rates of change was obtained only by giving increased weight to slowly evolving characters (e.g., transversion parsimony, weighted parsimony). UPGMA performed well only when branch lengths were close in length.