Relative Efficiencies of Simple and Complex Substitution Models in Estimating Divergence Times in Phylogenomics

Relative Efficiencies of Simple and Complex Substitution Models in Estimating Divergence Times in Phylogenomics
复制标题

DOI:
10.1093/molbev/msaa049
复制
发表时间:
2020-02
影响因子:
10.7
通讯作者:
Qiqing Tao;Jose Barba-Montoya;Louise A. Huuki;Mary Kathleen Durnan;Sudhir Kumar
Qiqing Tao;Jose Barba-Montoya;Louise A. Huuki;Mary Kathleen Durnan;Sudhir Kumar
中科院分区:
生物学1区
文献类型:
--
作者:
Qiqing Tao;Jose Barba-Montoya;Louise A. Huuki;Mary Kathleen Durnan;Sudhir Kumar

文献摘要

被引文献

相似文献

分子进化中的传统智慧是应用核苷酸和氨基酸取代的参数丰富的模型来估计分歧时间。然而,对于经常包含来自许多物种和基因的序列的当代数据集,高度复杂的模型与简单模型产生的时间估计值之间的差异的实际程度还有待量化。在重新分析的许多大型多物种的路线,从不同的类群使用相同的树拓扑结构和校准,我们发现,使用最简单的模型可以产生分歧的时间估计和可信区间类似的复杂的模型应用在原来的研究。这一结果令人惊讶,因为使用简单模型低估了所有分析数据集的序列差异。我们发现了三个基本原因,在许多实际的数据集的时间估计模型的复杂性观察到的鲁棒性。首先,最简单的模型下的分支长度和节点到尖端距离的估计值与使用最复杂的模型产生的估计值近似线性关系,特别是对于具有许多序列的数据集。第二,放松的时钟方法自动调整速率的分支,经历了相当大的低估序列的分歧,导致类似于那些从复杂的模型的时间估计。第三,即使在分析中包含一些良好的校准,也可以减少简单和复杂模型的时间估计差异。在这些经验数据分析中,时间估计对模型复杂性的鲁棒性是令人鼓舞的,因为所有的生物基因组学研究都使用统计模型,这些模型是对实际进化替代过程的过度简化描述。
The conventional wisdom in molecular evolution is to apply parameter-rich models of nucleotide and amino acid substitutions for estimating divergence times. However, the actual extent of the difference between time estimates produced by highly complex models compared to those from simple models is yet to be quantified for contemporary datasets that frequently contain sequences from many species and genes. In a reanalysis of many large multispecies alignments from diverse groups of taxa using the same tree topologies and calibrations, we found that the use of the simplest models can produce divergence time estimates and credibility intervals similar to those obtained from the complex models applied in the original studies. This result is surprising because the use of simple models underestimates sequence divergence for all the datasets analyzed. We find three fundamental reasons for the observed robustness of time estimates to model complexity in many practical datasets. First, the estimates of branch lengths and node-to-tip distances under the simplest model show an approximately linear relationship with those produced by using the most complex models applied, especially for datasets with many sequences. Second, relaxed clock methods automatically adjust rates on branches that experience considerable underestimation of sequence divergences, resulting in time estimates that are similar to those from complex models. And, third, the inclusion of even a few good calibrations in an analysis can reduce the difference in time estimates from simple and complex models. The robustness of time estimates to models complexity in these empirical data analyses is encouraging, because all phylogenomics studies use statistical models that are oversimplified descriptions of actual evolutionary substitution processes.