Adaptive MCMC in Bayesian phylogenetics: an application to analyzing partitioned data in BEAST

Adaptive MCMC in Bayesian phylogenetics: an application to analyzing partitioned data in BEAST
复制标题

DOI:
10.1093/bioinformatics/btx088
复制
发表时间:
2017-06-15
期刊:
影响因子:
5.8
通讯作者:
Suchard, Marc A.
Suchard, Marc A.
中科院分区:
生物学3区
文献类型:
--
作者:
Baele, Guy;Lemey, Philippe;Suchard, Marc A.

文献摘要

被引文献

相似文献

动机:测序技术的进步继续提供越来越大的分子序列数据集,这些数据集经常被严重分割,以便准确地模拟潜在的进化过程。在系统发育分析中,划分策略涉及估计不同基因和这些基因内不同位置的条件独立的分子进化模型,需要估计大量的进化参数,导致此类分析的计算负担增加。在过去的二十年里,我们也看到了多核处理器的兴起,无论是在中央处理器(CPU)还是图形处理单元处理器市场,都实现了大规模并行计算,而许多软件包还没有充分利用这些计算来进行多方分析。结果:我们在此提出了一种马尔可夫链蒙特卡罗(MCMC)方法,该方法使用自适应多变量转移核来并行估计大量参数,并利用多核处理在分区数据上进行分割。通过几个现实世界的例子,我们证明了我们的方法能够比通常使用单变量转换核的混合的标准方法更有效地估计这些多部参数。在一种情况下,在估计异构数据集中非编码分区的相对速率参数时,MCMC集成效率提高了>14倍。
Motivation: Advances in sequencing technology continue to deliver increasingly large molecular sequence datasets that are often heavily partitioned in order to accurately model the underlying evolutionary processes. In phylogenetic analyses, partitioning strategies involve estimating conditionally independent models of molecular evolution for different genes and different positions within those genes, requiring a large number of evolutionary parameters that have to be estimated, leading to an increased computational burden for such analyses. The past two decades have also seen the rise of multi-core processors, both in the central processing unit (CPU) and Graphics processing unit processor markets, enabling massively parallel computations that are not yet fully exploited by many software packages for multipartite analyses.Results: We here propose a Markov chain Monte Carlo (MCMC) approach using an adaptive multi-variate transition kernel to estimate in parallel a large number of parameters, split across partitioned data, by exploiting multi-core processing. Across several real-world examples, we demonstrate that our approach enables the estimation of these multipartite parameters more efficiently than standard approaches that typically use a mixture of univariate transition kernels. In one case, when estimating the relative rate parameter of the non-coding partition in a heterochronous dataset, MCMC integration efficiency improves by>14-fold.