An empirical examination of the utility of codon-substitution models in phylogeny reconstruction

An empirical examination of the utility of codon-substitution models in phylogeny reconstruction
复制标题

DOI:
10.1080/10635150500354688
复制
发表时间:
2005-10-01
期刊:
影响因子:
6.5
通讯作者:
Yang, ZH
Yang, ZH
中科院分区:
生物学1区
文献类型:
--
作者:
Ren, FR;Tanaka, H;Yang, ZH

文献摘要

被引文献

相似文献

密码子替换模型通常被用来比较蛋白质编码的DNA序列,并且在检测作用于蛋白质的自然选择信号方面特别有效。它们在重建分子系统学和测定物种分化时间方面的效用还没有被探索。密码子模型自然适应同义和非同义替换,它们以非常不同的速度发生,可能分别为现代和古代的分歧提供信息。因此,密码子模型有望有效地利用蛋白质编码的DNA序列中的系统发育信息。在这里,我们将密码子模型应用于来自8个酵母物种的106个蛋白质编码基因,以使用最大似然法重建系统发育,并与基于核苷酸和氨基酸的分析进行比较。结果似乎证实了这一预期。在简单的替代模型下,基于核苷酸的分析在恢复最近的差异方面是有效的,而基于氨基酸的分析在恢复深度差异方面表现得更好。密码子模型似乎结合了氨基酸和核苷酸数据的优势,在恢复最近和深度差异方面具有良好的性能。用氨基酸和密码子模型估计相对物种分化时间表明,将基因序列翻译成蛋白质导致信息损失从深层节点的30%到最近节点的66%。虽然计算负担使密码子模型不适用于大数据集中的树搜索,但我们建议它们可能有助于比较候选树。考虑到三个密码子位置的进化动力学差异的核苷酸模型也表现良好,计算成本要低得多。我们讨论了模型与数据的适合性与其在系统发育重建中的实用性之间的关系,并告诫不要使用过于复杂的替代模型。
Models of codon substitution have been commonly used to compare protein-coding DNA sequences and are particularly effective in detecting signals of natural selection acting on the protein. Their utility in reconstructing molecular phylogenies and in dating species divergences has not been explored. Codon models naturally accommodate synonymous and nonsynonymous substitutions, which occur at very different rates and may be informative for recent and ancient divergences, respectively. Thus codon models may be expected to make an efficient use of phylogenetic information in protein-coding DNA sequences. Here we applied codon models to 106 protein-coding genes from eight yeast species to reconstruct phylogenies using the maximum likelihood method, in comparison with nucleotide- and amino acid-based analyses. The results appeared to confirm that expectation. Nucleotide-based analysis, under simplistic substitution models, were efficient in recovering recent divergences whereas amino acid-based analysis performed better at recovering deep divergences. Codon models appeared to combine the advantages of amino acid and nucleotide data and had good performance at recovering both recent and deep divergences. Estimation of relative species divergence times using amino acid and codon models suggested that translation of gene sequences into proteins led to information loss of from 30% for deep nodes to 66% for recent nodes. Although computational burden makes codon models unfeasible for tree search in large data sets, we suggest that they may be useful for comparing candidate trees. Nucleotide models that accommodate the differences in evolutionary dynamics at the three codon positions also performed well, at much less computational cost. We discuss the relationship between a model's fit to data and its utility in phylogeny reconstruction and caution against use of overly complex substitution models.