The Implications of Interrelated Assumptions on Estimates of Divergence Times and Rates of Diversification

The Implications of Interrelated Assumptions on Estimates of Divergence Times and Rates of Diversification
复制标题

DOI:
10.1093/sysbio/syab021
复制
发表时间:
2021-03-24
期刊:
影响因子:
6.5
通讯作者:
Scotland, Robert W.
Scotland, Robert W.
中科院分区:
生物学1区
文献类型:
--
作者:
Carruthers, Tom;Scotland, Robert W.

文献摘要

被引文献

相似文献

系统发育正越来越多地被用作深入了解宏观进化史的基础。在这里,我们使用模拟实验和实证分析,以评估方法,使用遗传学作为基础,使发散时间和多样化率的估计。这是第一个研究提出了一个全面的评估的关键变量,在这一领域的分析,包括替代率,物种形成率,灭绝,加上字符采样和分类单元采样。我们表明,在不切实际的简单化的情况下(替代率和物种形成率是恒定的,并没有灭绝),增加字符和分类单元的采样导致更准确和精确的参数估计。相比之下,在更复杂但更现实的情况下(替代率、物种形成率和灭绝率各不相同),增加性状和分类单元取样的准确性和精确度的提高要有限得多。缺乏准确性和精确性甚至发生在使用旨在解释更复杂情况的方法时,例如放松时钟,化石校准以及允许物种形成率和灭绝率变化的模型。在分析基因组规模数据集时,问题也仍然存在。这些结果表明,当生成数据的过程更加复杂时,会出现两个相互关联的问题。首先,方法论假设更容易被违反。第二,数据信息内容的局限性变得更加重要。
Phylogenies are increasingly being used as a basis to provide insight into macroevolutionary history. Here, we use simulation experiments and empirical analyses to evaluate methods that use phylogenies as a basis to make estimates of divergence times and rates of diversification. This is the first study to present a comprehensive assessment of the key variables that underpin analyses in this field-including substitution rates, speciation rates, and extinction, plus character sampling and taxon sampling. We show that in unrealistically simplistic cases (where substitution rates and speciation rates are constant, and where there is no extinction), increased character and taxon sampling lead to more accurate and precise parameter estimates. By contrast, in more complex but realistic cases (where substitution rates, speciation rates, and extinction rates vary), gains in accuracy and precision from increased character and taxon sampling are far more limited. The lack of accuracy and precision even occurs when using methods that are designed to account for more complex cases, such as relaxed clocks, fossil calibrations, and models that allow speciation rates and extinction rates to vary. The problem also persists when analyzing genomic scale data sets. These results suggest two interrelated problems that occur when the processes that generated the data are more complex. First, methodological assumptions are more likely to be violated. Second, limitations in the information content of the data become more important.