Recoding Amino Acids to a Reduced Alphabet may Increase or Decrease Phylogenetic Accuracy

Recoding Amino Acids to a Reduced Alphabet may Increase or Decrease Phylogenetic Accuracy
复制标题

DOI:
10.1093/sysbio/syac042
复制
发表时间:
2022-07-28
期刊:
影响因子:
6.5
通讯作者:
Embley, T. Martin
Embley, T. Martin
中科院分区:
生物学1区
文献类型:
--
作者:
Foster, Peter G.;Schrempf, Dominik;Embley, T. Martin

文献摘要

被引文献

相似文献

在使用氨基酸数据时,常见的分子系统发育特征,如长分支和成分异质性,可能会对系统发育重建造成问题。在系统发育分析之前,将比对重新编码为简化的字母表通常被用来探索和潜在地减少此类问题的影响。我们使用四分类树上的模拟数据来检验该策略在拓扑精度上的有效性。我们以具有系统发育挑战性的方式模拟比对,以测试分析的系统发育准确性,使用各种重新编码策略和常用的同质模型。我们测试了三种基于氨基酸互换性的重新编码方法,以及另一种基于降低比对序列之间组成异质性的重新编码方法,该方法通过卡方统计来衡量。我们的模拟结果表明,在序列接近饱和的长枝树上,基于可交换性的编码对精度影响不大,但基于卡方的编码降低了精度。然后,我们模拟了树上具有不同种类组成异质性的序列。重新编码通常会提高此类比对的准确性。基于可交换性的重新编码很少比不重新编码更糟糕,而且往往要好得多。基于降低卡方差值的重新编码在某些情况下提高了准确性,但在另一些情况下却没有,这表明低成分异质性本身不足以提高这些比对分析的准确性。我们还使用特定位置的氨基酸图谱来模拟比对,使序列的组成在比对位置上具有异质性。对于这些数据集,基于交换性的重新编码与站点同质模型相结合的准确性较差,但基于卡方的重新编码提高了准确性。然后,我们模拟了站点和树的组成都是异质的数据集,就像许多真实的数据集一样。根据组成树异质性的类型和重新编码方案的不同,对重新编码这种双重问题数据集的准确性的影响差别很大。有趣的是,用NDCH或CAT模型分析未记录的成分异质性比对通常比同质分析更准确,无论是否记录。总体而言,我们的结果表明,为重新编码的氨基酸数据集制作树可能是有用的,但需要谨慎地解释为更全面分析的一部分。使用NDCH和CAT等更好的模型,可以直接解释数据中的模式,可能会为分析经验数据提供更有前途的长期解决方案。[成分异质性;进化模型;系统发育方法;重新编码氨基酸数据集。]
Common molecular phylogenetic characteristics such as long branches and compositional heterogeneity can be problematic for phylogenetic reconstruction when using amino acid data. Recoding alignments to reduced alphabets before phylogenetic analysis has often been used both to explore and potentially decrease the effect of such problems. We tested the effectiveness of this strategy on topological accuracy using simulated data on four-taxon trees. We simulated alignments in phylogenetically challenging ways to test the phylogenetic accuracy of analyses using various recoding strategies together with commonly used homogeneous models. We tested three recoding methods based on amino acid exchangeability, and another recoding method based on lowering the compositional heterogeneity among alignment sequences as measured by the Chi-squared statistic. Our simulation results show that on trees with long branches where sequences approach saturation, accuracy was not greatly affected by exchangeability-based recodings, but Chi-squared-based recoding decreased accuracy. We then simulated sequences with different kinds of compositional heterogeneity over the tree. Recoding often increased accuracy on such alignments. Exchangeability-based recoding was rarely worse than not recoding, and often considerably better. Recoding based on lowering the Chi-squared value improved accuracy in some cases but not in others, suggesting that low compositional heterogeneity by itself is not sufficient to increase accuracy in the analysis of these alignments. We also simulated alignments using site-specific amino acid profiles, making sequences that had compositional heterogeneity over alignment sites. Exchangeability-based recoding coupled with site-homogeneous models had poor accuracy for these data sets but Chi-squared-based recoding on these alignments increased accuracy. We then simulated data sets that were compositionally both site- and tree-heterogeneous, like many real data sets. The effect on the accuracy of recoding such doubly problematic data sets varied widely, depending on the type of compositional tree heterogeneity and on the recoding scheme. Interestingly, analysis of unrecoded compositionally heterogeneous alignments with the NDCH or CAT models was generally more accurate than homogeneous analysis, whether recoded or not. Overall, our results suggest that making trees for recoded amino acid data sets can be useful, but they need to be interpreted cautiously as part of a more comprehensive analysis. The use of better-fitting models like NDCH and CAT, which directly account for the patterns in the data, may offer a more promising long-term solution for analyzing empirical data. [Compositional heterogeneity; models of evolution; phylogenetic methods; recoding amino acid data sets.]