MAXIMUM-LIKELIHOOD PHYLOGENETIC ESTIMATION FROM DNA-SEQUENCES WITH VARIABLE RATES OVER SITES - APPROXIMATE METHODS

MAXIMUM-LIKELIHOOD PHYLOGENETIC ESTIMATION FROM DNA-SEQUENCES WITH VARIABLE RATES OVER SITES - APPROXIMATE METHODS
复制标题

DOI:
10.1007/bf00160154
复制
发表时间:
1994-09-01
影响因子:
3.9
通讯作者:
YANG, ZH
YANG, ZH
中科院分区:
生物学3区
文献类型:
--
作者:
YANG, ZH

文献摘要

被引文献

相似文献

提出了两种近似方法用于最大似然系统发育估计,其允许跨核苷酸位点的可变替换率。对具有完全不同特征的三个数据集进行了分析,以实证检验这些方法的性能。第一个称为“离散伽玛模型”,使用多个类别的速率来近似伽玛分布,每个类别的概率相等。每个类别的平均值用于表示属于该类别的所有比率。发现该方法的性能非常好,并且四个这样的类别似乎足以产生模型与数据的最佳或接近最佳拟合,以及对连续分布的可接受的近似。第二种方法称为“固定费率模型”,根据假设星树的预测费率将站点分为几个类别。当评估其他树拓扑时,假设不同类别的站点以这些固定速率发展。对数据集的分析表明,该方法可以产生合理的结果,但它似乎具有最小二乘成对比较的一些属性;例如,非最佳树的内部分支长度通常为零。这两种方法的计算要求与 Felsenstein(1981,J Mol Evol 17:368-376)模型的计算要求相当,该模型假设所有位点采用单一速率。
Two approximate methods are proposed for maximum likelihood phylogenetic estimation, which allow variable rates of substitution across nucleotide sites. Three data sets with quite different characteristics were analyzed to examine empirically the performance of these methods. The first, called the ''discrete gamma model,'' uses several categories of rates to approximate the gamma distribution, with equal probability for each category. The mean of each category is used to represent all the rates falling in the category. The performance of this method is found to be quite good, and four such categories appear to be sufficient to produce both an optimum, or near-optimum fit by the model to the data, and also an acceptable approximation to the continuous distribution. The second method, called ''fixed-rates model,'' classifies sites into several classes according to their rates predicted assuming the star tree. Sites in different classes are then assumed to be evolving at these fixed rates when other tree topologies are evaluated. Analyses of the data sets suggest that this method can produce reasonable results, but it seems to share some properties of a least-squares pairwise comparison; for example, interior branch lengths in nonbest trees are often found to be zero. The computational requirements of the two methods are comparable to that of Felsenstein's (1981, J Mol Evol 17:368-376) model, which assumes a single rate for all the sites.