Incorporating Hierarchical Characters into Phylogenetic Analysis

Incorporating Hierarchical Characters into Phylogenetic Analysis
复制标题

DOI:
10.1093/sysbio/syab005
复制
发表时间:
2021-02-09
期刊:
影响因子:
6.5
通讯作者:
St John, Katherine
St John, Katherine
中科院分区:
生物学1区
文献类型:
--
作者:
Hopkins, Melanie J.;St John, Katherine

文献摘要

被引文献

相似文献

流行的系统发育树的最优标准集中在适用于所有分类群的字符序列。随着研究范围的扩大,可能会出现某些性状适用于部分分类群而不适用于其他分类群的情况。过去的工作已经探讨了处理不适用的字符作为缺失数据的局限性,注意到这种策略可能有利于树内部节点被分配不可能的状态,在亚分支内的类群的安排是不适当的影响,在遥远的部分的树的变化,和/或在分类群,否则共享大多数主要字符远分组。最近已经提出了避免前两个问题的方法。在这里,我们提出了一种替代方法,避免所有三个问题。我们专注于数据矩阵,使用还原编码的特点,也就是说,明确纳入先天层次结构诱导的不适用性,因此,我们的方法扩展到层次结构的字符,一般。在最大简约的精神,所提出的标准寻求系统发育树的任何树分支的变化最小,但其中的变化是定义在不同的度量,衡量不适用的字符的影响。该方法可以容纳二进制,多态,有序,无序和多态字符。受Fitch算法的启发,我们给出了一个多项式时间的算法,用于在一系列相异度量下对树进行评分,并证明了算法的正确性。我们表明,由此产生的最优性标准是计算困难的,减少到NP-硬度的最大简约最优性标准。我们证明了我们的方法使用合成和经验数据集,并与其他最近提出的方法选择最佳的系统发育树时,数据包括层次特征的结果进行比较。
Popular optimality criteria for phylogenetic trees focus on sequences of characters that are applicable to all the taxa. As studies grow in breadth, it can be the case that some characters are applicable for a portion of the taxa and inapplicable for others. Past work has explored the limitations of treating inapplicable characters as missing data, noting that this strategy may favor trees where internal nodes are assigned impossible states, where the arrangement of taxa within subclades is unduly influenced by variation in distant parts of the tree, and/or where taxa that otherwise share most primary characters are grouped distantly. Approaches that avoid the first two problems have recently been proposed. Here, we propose an alternative approach which avoids all three problems. We focus on data matrices that use reductive coding of traits, that is, explicitly incorporate the innate hierarchy induced by inapplicability, and as such our approach extend to hierarchical characters, in general. In the spirit of maximum parsimony, the proposed criterion seeks the phylogenetic tree with the minimal changes across any tree branch, but where changes are defined in terms of dissimilarity metrics that weigh the effects of inapplicable characters. The approach can accommodate binary, multistate, ordered, unordered, and polymorphic characters. We give a polynomial-time algorithm, inspired by Fitch's algorithm, to score trees under a family of dissimilarity metrics, and prove its correctness. We show that the resulting optimality criteria is computationally hard, by reduction to the NP-hardness of the maximum parsimony optimality criteria. We demonstrate our approach using synthetic and empirical data sets and compare the results with other recently proposed methods for choosing optimal phylogenetic trees when the data includes hierarchical characters.