Statistically consistent and computationally efficient inference of ancestral DNA sequences in the TKF91 model under dense taxon sampling

Statistically consistent and computationally efficient inference of ancestral DNA sequences in the TKF91 model under dense taxon sampling
复制标题

在密集分类单元采样下,TKF91 模型中祖先 DNA 序列的统计一致性和计算效率推断

DOI:
10.1007/s11538-020-00693-3
复制
发表时间:
2020
影响因子:
3.5
通讯作者:
Roch, Sebastien
Roch, Sebastien
中科院分区:
数学4区
文献类型:
--
作者:
Fan, Wai-Tong;Roch, Sebastien

文献摘要

参考文献

被引文献

相似文献

在进化生物学中,生命有机体的物种形成历史由系统发育图来表示,即有根的树,其叶子与当前物种相对应,其分枝表明过去的物种形成事件。系统发育分析通常依赖于从感兴趣的物种收集的分子序列,例如DNA序列,在这种情况下,通常使用基于树上序列进化的随机模型的统计方法。对于易操纵性,这些模型必须对所涉及的进化机制做出简化的假设。特别是,通常被省略的是核苷酸的插入和缺失--也被称为INDELs。在统计系统发育分析中适当地考虑Indels仍然是计算进化生物学的一个主要挑战。在这里,我们考虑在一个包含核苷酸替换、插入和缺失的序列进化模型中重建已知系统发育上的祖先序列的问题,特别是经典的TKF91过程。我们专注于有界高度的密集系统发育的情况,我们称之为富含分类单元的环境,在那里统计一致性是可以实现的。我们给出了第一个在恒定突变率下具有可证明保证的显式重建算法。当系统发育满足“大爆炸”条件时,我们的算法成功,“大爆炸”条件是在这种情况下统计一致性的一个充要条件。
In evolutionary biology, the speciation history of living organisms is represented graphically by a phylogeny, that is, a rooted tree whose leaves correspond to current species and whose branchings indicate past speciation events. Phylogenetic analyses often rely on molecular sequences, such as DNA sequences, collected from the species of interest, and it is common in this context to employ statistical approaches based on stochastic models of sequence evolution on a tree. For tractability, such models necessarily make simplifying assumptions about the evolutionary mechanisms involved. In particular, commonly omitted areinsertionsanddeletionsof nucleotides—also known as indels. Properly accounting for indels in statistical phylogenetic analyses remains a major challenge in computational evolutionary biology. Here, we consider the problem of reconstructing ancestral sequences on a known phylogeny in a model of sequence evolution incorporating nucleotide substitutions, insertions and deletions, specifically the classical TKF91 process. We focus on the case of dense phylogenies of bounded height, which we refer to as the taxon-rich setting, where statistical consistency is achievable. We give the first explicit reconstruction algorithm with provable guarantees under constant rates of mutation. Our algorithm succeeds when the phylogeny satisfies the “big bang” condition, a necessary and sufficient condition for statistical consistency in this setting.
DOI: 10.1214/18-ejp165
发表时间: 2018
影响因子: 1.4
作者:
Fan, Wai-Tong;Roch, Sebastien
通讯作者: Roch, Sebastien
DOI: 10.1214/aoap/998926994
发表时间: 2001
影响因子: 1.8
作者:
Elchanan Mossel
通讯作者: Elchanan Mossel
DOI: 10.1016/j.spa.2012.08.004
发表时间: 2009-12
期刊: ArXiv
影响因子: --
作者:
Alexandr Andoni;C. Daskalakis;Avinatan Hassidim;S. Roch
通讯作者: Alexandr Andoni;C. Daskalakis;Avinatan Hassidim;S. Roch
DOI: 10.1007/bf02193625
发表时间: 1991-08-01
影响因子: 3.9
作者:
THORNE, JL;KISHINO, H;FELSENSTEIN, J
通讯作者: FELSENSTEIN, J
DOI: 10.1007/978-3-540-69903-3_1
发表时间: 2008-07
期刊: --
影响因子: --
作者:
M. Mitzenmacher
通讯作者: M. Mitzenmacher