Handling Logical Character Dependency in Phylogenetic Inference: Extensive Performance Testing of Assumptions and Solutions Using Simulated and Empirical Data

Handling Logical Character Dependency in Phylogenetic Inference: Extensive Performance Testing of Assumptions and Solutions Using Simulated and Empirical Data
复制标题

处理系统发育推断中的逻辑字符依赖性:使用模拟和经验数据对假设和解决方案进行广泛的性能测试

DOI:
10.1093/sysbio/syad006
复制
发表时间:
2023
期刊:
影响因子:
6.5
通讯作者:
Wright, April M
Wright, April M
中科院分区:
生物学1区
文献类型:
--
作者:
Simões, Tiago R;Vernygora, Oksana V;de Medeiros, Bruno A;Wright, April M

文献摘要

相似文献

在形态学数据集的系统发育推断中,逻辑字符依赖是一个主要的概念和方法问题,因为它违反了所有系统发育方法所共有的字符独立性假设。它更常见于更高层次的遗传学或表征主要进化转变的数据集中,因为这些代表了生命树的部分,(主要)解剖特征要么起源要么完全消失。因此,与这些主要性状相关的次要性状在所有不存在该性状的样本分类群中变得“不适用”。在过去的三十年里,已经探索了各种解决方案来处理字符依赖性,例如替代字符编码方案以及最近的新算法实现。然而,所提出的解决方案的准确性,或字符依赖性在不同的最优性标准的影响,从来没有直接使用标准的性能指标进行测试。在这里,我们利用简单和复杂的模拟形态数据集分析不同的最大简约优化程序和贝叶斯推理测试字符依赖的各种编码和算法解决方案的准确性。这是补充的经验分析,使用重新编码的数据集古颚类鸟类。我们发现,在小的,模拟的数据集,缺席编码比其他流行的编码策略(偶然和多态),而在更复杂的模拟(较大的数据集控制不同的树结构和字符分布模型)偶然编码更频繁地青睐。在偶然编码下,最近提出的加权算法产生最大简约最准确的结果。然而,贝叶斯推理优于所有基于简约的解决方案,以处理字符依赖性,由于其优化过程的根本差异,一个简单的替代方案,一直被忽视。然而,我们表明,更多的主要特征轴承次要(依赖)性状有在一个数据集,更难估计真正的系统发育树,无论最优标准,由于相当大的扩展树参数空间。[贝叶斯推理,字符依赖性,字符编码,距离度量,形态遗传学,最大简约性,性能,系统发育准确性。]
Logical character dependency is a major conceptual and methodological problem in phylogenetic inference of morphological data sets, as it violates the assumption of character independence that is common to all phylogenetic methods. It is more frequently observed in higher-level phylogenies or in data sets characterizing major evolutionary transitions, as these represent parts of the tree of life where (primary) anatomical characters either originate or disappear entirely. As a result, secondary traits related to these primary characters become “inapplicable” across all sampled taxa in which that character is absent. Various solutions have been explored over the last three decades to handle character dependency, such as alternative character coding schemes and, more recently, new algorithmic implementations. However, the accuracy of the proposed solutions, or the impact of character dependency across distinct optimality criteria, has never been directly tested using standard performance measures. Here, we utilize simple and complex simulated morphological data sets analyzed under different maximum parsimony optimization procedures and Bayesian inference to test the accuracy of various coding and algorithmic solutions to character dependency. This is complemented by empirical analyses using a recoded data set on palaeognathid birds. We find that in small, simulated data sets, absent coding performs better than other popular coding strategies available (contingent and multistate), whereas in more complex simulations (larger data sets controlled for different tree structure and character distribution models) contingent coding is favored more frequently. Under contingent coding, a recently proposed weighting algorithm produces the most accurate results for maximum parsimony. However, Bayesian inference outperforms all parsimony-based solutions to handle character dependency due to fundamental differences in their optimization procedures—a simple alternative that has been long overlooked. Yet, we show that the more primary characters bearing secondary (dependent) traits there are in a data set, the harder it is to estimate the true phylogenetic tree, regardless of the optimality criterion, owing to a considerable expansion of the tree parameter space. [Bayesian inference, character dependency, character coding, distance metrics, morphological phylogenetics, maximum parsimony, performance, phylogenetic accuracy.]