Ancestral state reconstruction with large numbers of sequences and edge-length estimation

Ancestral state reconstruction with large numbers of sequences and edge-length estimation
复制标题

使用大量序列和边长估计进行祖先状态重建

DOI:
10.1007/s00285-022-01715-5
复制
发表时间:
2021
影响因子:
1.9
通讯作者:
E. Susko
E. Susko
中科院分区:
数学4区
文献类型:
--
作者:
L. Ho;E. Susko

文献摘要

参考文献

被引文献

相似文献

基于似然的方法被广泛认为是重建祖先状态的最佳方法。尽管研究这些方法的性质已经付出了很多努力,但以前的工作通常假设树的拓扑结构和边长度都是已知的。在某些情况下,所研究的分类群可能对树的拓扑结构相当熟悉。然而,当序列长度远远小于物种数量时,边缘长度不可能被准确估计。我们研究了在星形树下离散特征祖先状态的极大似然估计和经验贝叶斯估计的一致性。我们证明了基于似然的重构在对称模型下是一致的,而在非对称模型下是不一致的。然而,我们证明了在非对称模型下祖先状态的简单一致估计是可用的。结果表明,当考虑的序列数量变得非常大时,似然方法可能出乎意料地具有不希望的性质。讨论了结果的更广泛含义。
Likelihood-based methods are widely considered the best approaches for reconstructing ancestral states. Although much effort has been made to study properties of these methods, previous works often assume that both the tree topology and edge lengths are known. In some scenarios the tree topology might be reasonably well known for the taxa under study. When sequence length is much smaller than the number of species, however, edge lengths are not likely to be accurately estimated. We study the consistency of the maximum likelihood and empirical Bayes estimators of the ancestral state of discrete traits in such settings under a star tree. We prove that the likelihood-based reconstruction is consistent under symmetric models but can be inconsistent under non-symmetric models. We show, however, that a simple consistent estimator for the ancestral states is available under non-symmetric models. The results illustrate that likelihood methods can unexpectedly have undesirable properties as the number of sequences considered gets very large. Broader implications of the results are discussed.
DOI: 10.1214/18-ejp165
发表时间: 2018
影响因子: 1.4
作者:
Fan, Wai-Tong;Roch, Sebastien
通讯作者: Roch, Sebastien