Accuracies of ancestral amino acid sequences inferred by the parsimony, likelihood, and distance methods

Accuracies of ancestral amino acid sequences inferred by the parsimony, likelihood, and distance methods
复制标题

DOI:
10.1007/pl00000067
复制
发表时间:
1997-01-01
影响因子:
3.9
通讯作者:
Nei, M
Nei, M
中科院分区:
生物学3区
文献类型:
--
作者:
Zhang, JZ;Nei, M

文献摘要

被引文献

相似文献

有关祖先生物蛋白质序列的信息对于识别引起蛋白质进化中功能变化的关键氨基酸替换是重要的。利用计算机模拟,我们研究了祖先氨基酸推断的准确性,目前可用的两种方法(最大简约[MP]和最大似然[ML]方法),除了距离的方法,这是本文新开发的。这三种方法在氨基酸序列差异较小的情况下都能给出可靠的推断。然而,当序列分歧程度高时,ML和距离方法比MP方法给出更准确的结果,特别是当系统发育树包括长分支时。当添加或删除一些现今的序列时,推断的祖先氨基酸的准确性不会发生很大变化。当ML和距离方法使用不正确的氨基酸取代模型时,准确度降低,但仍高于MP方法。当使用的树拓扑部分不正确时,树的正确部分的准确性实际上不受影响。通过ML和距离方法计算的推断的祖先氨基酸的后验概率是使用正确的取代模型时真实概率的无偏估计,但当使用更简单的模型时可能会被高估。
Information about protein sequences of ancestral organisms is important for identifying critical amino acid substitutions that have caused the functional change of proteins in evolution. Using computer simulation, we studied the accuracy of ancestral amino acids inferred by two currently available methods (maximum-parsimony [MP] and maximum-likelihood [ML] methods) in addition to a distance method, which was newly developed in this paper. All three methods give reliable inference when the divergence of amino acid sequences is low. When the extent of sequence divergence is high, however, the ML and distance methods give more accurate results than the MP method, particularly when the phylogenetic tree includes long branches. The accuracy of inferred ancestral amino acids does not change very much when a few present-day sequences are added or eliminated. When an incorrect model of amino acid substitution is used for the ML and distance methods, the accuracy decreases, but it is still higher than that for the MP method. When the tree topology used is partially incorrect, the accuracy in the correct part of the tree is virtually unaffected. The posterior probability of inferred ancestral amino acids computed by the ML and distance methods is an unbiased estimate of the true probability when a correct substitution model is used but may become an overestimate when a simpler model is used.