Bayesian inference of sampled ancestor trees for epidemiology and fossil calibration.

Bayesian inference of sampled ancestor trees for epidemiology and fossil calibration.
复制标题

采样祖先树的贝叶斯推断用于流行病学和化石校准。

DOI:
10.1371/journal.pcbi.1003919
复制
发表时间:
2014-12
影响因子:
4.3
通讯作者:
Drummond AJ
Drummond AJ
中科院分区:
生物学2区
文献类型:
--
作者:
Gavryushkina A;Welch D;Stadler T;Drummond AJ

文献摘要

参考文献

被引文献

相似文献

系统发育分析,包括化石或分子序列的采样,需要模型允许一个样本是另一个样本的直接祖先。由于以前可用的系统发育推断工具假设所有样本都是提示,因此它们不允许这种可能性。我们已经开发并实现了一个贝叶斯马尔可夫链蒙特卡罗(MCMC)算法来推断我们所谓的采样祖先树,即采样个体可以是其他采样个体的直接祖先的树。我们使用了一组出生-死亡模型,其中个体在采样后可能仍处于树形过程中,特别是我们将出生-死亡天际线模型(Stadler et al., 2013)扩展到采样的祖先树。这种方法允许检测采样的祖先,以及估计个体在采样时将从过程中移除的概率。我们表明,即使采样的祖先在分析中不是特别感兴趣,未能考虑到它们会导致参数估计中的显着偏差。我们还表明,每个样本来自不同时间点的采样祖先出生-死亡模型是不可识别的,因此需要知道一个参数才能推断其他参数。我们将祖先样本的系统发育推断应用于流行病学数据,其中祖先样本的可能性使我们能够识别采样后感染其他个体的个体,并推断基本的流行病学参数。当化石与现存物种样本一起包含时,我们也应用该方法来推断分化时间和多样化率,因此化石事件被建模为树分支过程的一部分。如文献所述,这种建模有许多优点。采样器可以作为开源BEAST2包获得(https://github.com/CompEvol/sampled-ancestors)。系统发育分析的核心目标是从分子数据中估计进化关系和进化分支过程的动态参数(如宏观进化或流行病学参数)。在这些分析中使用的统计方法要求指定潜在的树分支过程。分支过程的标准模型最初是为了描述现代物种的进化历史而设计的,它不允许一个样本分类单元是另一个样本分类单元的祖先。然而,对于许多类型的数据,直接祖先抽样的概率是不可忽略的。例如,当化石和现存物种一起分析以推断物种分化时间时,化石物种可能是也可能不是现存物种的直系祖先。在流行病学中,采样个体(从中获得病原体序列的宿主)可以在采样后感染其他个体,然后这些个体继续对自己进行采样。考虑直系祖先的模型产生的系统发育树与传统的系统发育树结构不同,因此在推理中使用这些模型需要新的计算方法。在这里,我们开发了一种系统发育分析方法,可以解释直系祖先的可能性。
Phylogenetic analyses which include fossils or molecular sequences that are sampled through time require models that allow one sample to be a direct ancestor of another sample. As previously available phylogenetic inference tools assume that all samples are tips, they do not allow for this possibility. We have developed and implemented a Bayesian Markov Chain Monte Carlo (MCMC) algorithm to infer what we call sampled ancestor trees, that is, trees in which sampled individuals can be direct ancestors of other sampled individuals. We use a family of birth-death models where individuals may remain in the tree process after sampling, in particular we extend the birth-death skyline model [Stadler et al., 2013] to sampled ancestor trees. This method allows the detection of sampled ancestors as well as estimation of the probability that an individual will be removed from the process when it is sampled. We show that even if sampled ancestors are not of specific interest in an analysis, failing to account for them leads to significant bias in parameter estimates. We also show that sampled ancestor birth-death models where every sample comes from a different time point are non-identifiable and thus require one parameter to be known in order to infer other parameters. We apply our phylogenetic inference accounting for sampled ancestors to epidemiological data, where the possibility of sampled ancestors enables us to identify individuals that infected other individuals after being sampled and to infer fundamental epidemiological parameters. We also apply the method to infer divergence times and diversification rates when fossils are included along with extant species samples, so that fossilisation events are modelled as a part of the tree branching process. Such modelling has many advantages as argued in the literature. The sampler is available as an open-source BEAST2 package (https://github.com/CompEvol/sampled-ancestors). A central goal of phylogenetic analysis is to estimate evolutionary relationships and the dynamical parameters underlying the evolutionary branching process (e.g. macroevolutionary or epidemiological parameters) from molecular data. The statistical methods used in these analyses require that the underlying tree branching process is specified. Standard models for the branching process which were originally designed to describe the evolutionary past of present day species do not allow one sampled taxon to be the ancestor of another. However the probability of sampling a direct ancestor is not negligible for many types of data. For example, when fossil and living species are analysed together to infer species divergence times, fossil species may or may not be direct ancestors of living species. In epidemiology, a sampled individual (a host from which a pathogen sequence was obtained) can infect other individuals after sampling, which then go on to be sampled themselves. The models that account for direct ancestors produce phylogenetic trees with a different structure from classic phylogenetic trees and so using these models in inference requires new computational methods. Here we developed a method for phylogenetic analysis that accounts for the possibility of direct ancestors.
DOI: 10.1093/sysbio/sys058
发表时间: 2012-12-01
期刊: Systematic biology
影响因子: 6.5
作者:
Ronquist F;Klopfstein S;Vilhelmsen L;Schulmeister S;Murray DL;Rasnitsyn AP
通讯作者: Rasnitsyn AP
DOI: 10.1080/106351501753462876
发表时间: 2001-11-01
期刊: SYSTEMATIC BIOLOGY
影响因子: 6.5
作者:
Lewis, PO
通讯作者: Lewis, PO
DOI: 10.1016/j.jtbi.2012.08.046
发表时间: 2012-12-21
影响因子: 2
作者:
Didier, Gilles;Royer-Carenzi, Manuela;Laurin, Michel
通讯作者: Laurin, Michel
DOI: 10.1371/journal.pbio.0040088
发表时间: 2006-05
期刊: PLoS biology
影响因子: 9.8
作者:
Drummond AJ;Ho SY;Phillips MJ;Rambaut A
通讯作者: Rambaut A
DOI: 10.1093/oxfordjournals.molbev.a025731
发表时间: 1997-12-01
影响因子: 10.7
作者:
Sanderson, MJ
通讯作者: Sanderson, MJ