A simulation approach for change-points on phylogenetic trees.

A simulation approach for change-points on phylogenetic trees.
复制标题

系统发育树上变化点的模拟方法。

DOI:
10.1089/cmb.2014.0218
复制
发表时间:
2015
期刊:
a journal of computational molecular cell biology
影响因子:
--
通讯作者:
Persing A
Persing A
中科院分区:
--
文献类型:
--
作者:
Persing A

文献摘要

相似文献

我们在每个ofmsites的序列,并假设他们已经从一个祖先的序列,形成一个已知的拓扑结构和分支长度的二叉树的根,但在内部节点的序列状态是未知的。树的拓扑结构和分支长度对于所有地点都是相同的,但进化模型的参数可能因地点而异。我们假设这些参数的分段常数模型,具有未知数量的变点,因此我们寻求执行贝叶斯推断的跨维参数空间。我们提出了两个新的想法来处理这种推理的计算挑战。首先,我们基于时间机器原理来近似模型:二叉树的顶部节点(靠近根)被真实分布的近似值所取代;随着更多的节点从树的顶部移除,计算似然的成本线性降低。这种方法引入了一种偏见,我们根据经验进行了调查。其次,我们开发了一个粒子边缘Metropolis-Hastings(PMMH)算法,它采用了顺序蒙特卡罗(SMC)采样器,可以使用第一个想法。我们的时间机器PMMH算法很好地应对标准计算算法的瓶颈之一:后验分布的跨维性质。该算法实现模拟和真实的数据的例子,我们经验证明其潜在的竞争力优于基于近似贝叶斯计算(ABC)技术的方法。
We observensequences at each ofmsites and assume that they have evolved from an ancestral sequence that forms the root of a binary tree of known topology and branch lengths, but the sequence states at internal nodes are unknown. The topology of the tree and branch lengths are the same for all sites, but the parameters of the evolutionary model can vary over sites. We assume a piecewise constant model for these parameters, with an unknown number of change-points and hence a transdimensional parameter space over which we seek to perform Bayesian inference. We propose two novel ideas to deal with the computational challenges of such inference. Firstly, we approximate the model based on the time machine principle: the top nodes of the binary tree (near the root) are replaced by an approximation of the true distribution; as more nodes are removed from the top of the tree, the cost of computing the likelihood is reduced linearly inn. The approach introduces a bias, which we investigate empirically. Secondly, we develop a particle marginal Metropolis-Hastings (PMMH) algorithm, that employs a sequential Monte Carlo (SMC) sampler and can use the first idea. Our time-machine PMMH algorithm copes well with one of the bottle-necks of standard computational algorithms: the transdimensional nature of the posterior distribution. The algorithm is implemented on simulated and real data examples, and we empirically demonstrate its potential to outperform competing methods based on approximate Bayesian computation (ABC) techniques.