DNA Sequences Are as Useful as Protein Sequences for Inferring Deep Phylogenies.

DNA Sequences Are as Useful as Protein Sequences for Inferring Deep Phylogenies.
复制标题

DOI:
10.1093/sysbio/syad036
复制
发表时间:
2023-11-01
期刊:
影响因子:
6.5
通讯作者:
--
中科院分区:
生物学1区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

深层系统发育的推断几乎完全使用蛋白质而不是DNA序列,因为蛋白质序列比DNA序列更不容易出现同质性和饱和或组成异质性问题。在这里,我们分析了一个理想遗传密码下的密码子进化模型,并证明这些观念可能是误解。我们进行了一项模拟研究,以评估蛋白质与DNA序列在推断深层系统发育方面的效用,在序列中跨位点和谱系之间的异质替代过程模型下生成蛋白质编码数据,然后使用核苷酸、氨基酸和密码子模型进行分析。在核苷酸取代模型下的DNA序列分析(可能排除了第三个密码子位置)恢复正确树的频率至少与在现代氨基酸模型下分析相应的蛋白质序列一样高。我们还将不同的数据分析策略应用于经验数据集,以推断后生动物的系统发育。我们的模拟和真实数据的结果表明,DNA序列在推断深层系统发育方面可能与蛋白质一样有用,不应被排除在此类分析之外。与蛋白质数据分析相比,在核苷酸模型下分析DNA数据具有主要的计算优势,这可能使得在推断深层系统发育时使用先进的模型来解释核苷酸取代过程中的位点间和谱系间异质性成为可能。
Inference of deep phylogenies has almost exclusively used protein rather than DNA sequences based on the perception that protein sequences are less prone to homoplasy and saturation or to issues of compositional heterogeneity than DNA sequences. Here, we analyze a model of codon evolution under an idealized genetic code and demonstrate that those perceptions may be misconceptions. We conduct a simulation study to assess the utility of protein versus DNA sequences for inferring deep phylogenies, with protein-coding data generated under models of heterogeneous substitution processes across sites in the sequence and among lineages on the tree, and then analyzed using nucleotide, amino acid, and codon models. Analysis of DNA sequences under nucleotide-substitution models (possibly with the third codon positions excluded) recovered the correct tree at least as often as analysis of the corresponding protein sequences under modern amino acid models. We also applied the different data-analysis strategies to an empirical dataset to infer the metazoan phylogeny. Our results from both simulated and real data suggest that DNA sequences may be as useful as proteins for inferring deep phylogenies and should not be excluded from such analyses. Analysis of DNA data under nucleotide models has a major computational advantage over protein-data analysis, potentially making it feasible to use advanced models that account for among-site and among-lineage heterogeneity in the nucleotide-substitution process in inference of deep phylogenies.
模棱两可的编码允许从汇总状态空间中的比对准确推断进化参数。
DOI: 10.1093/sysbio/syaa036
发表时间: 2021-01-01
期刊: Systematic biology
影响因子: 6.5
作者:
Weber CC;Perron U;Casey D;Yang Z;Goldman N
通讯作者: Goldman N
DOI: 10.1093/nar/gkq291
发表时间: 2010-07
影响因子: 14.9
作者:
Abascal F;Zardoya R;Telford MJ
通讯作者: Telford MJ
DOI: 10.1016/j.cub.2017.11.008
发表时间: 2017-12-18
期刊: CURRENT BIOLOGY
影响因子: 9.2
作者:
Feuda, Roberto;Dohrmann, Martin;Pisani, Davide
通讯作者: Pisani, Davide
DOI: 10.1093/molbev/msx281
发表时间: 2018-02-01
影响因子: 10.7
作者:
Hoang DT;Chernomor O;von Haeseler A;Minh BQ;Vinh LS
通讯作者: Vinh LS
DOI: 10.1093/molbev/msp098
发表时间: 2009-08
影响因子: 10.7
作者:
Fletcher W;Yang Z
通讯作者: Yang Z