Stochastic Variational Inference for Bayesian Phylogenetics: A Case of CAT Model

Stochastic Variational Inference for Bayesian Phylogenetics: A Case of CAT Model
复制标题

DOI:
10.1093/molbev/msz020
复制
发表时间:
2019-04-01
影响因子:
10.7
通讯作者:
Kishino, Hirohisa
Kishino, Hirohisa
中科院分区:
生物学1区
文献类型:
--
作者:
Dang, Tung;Kishino, Hirohisa

文献摘要

被引文献

相似文献

分子进化的模式在基因组中的不同基因位置和基因之间有所不同。通过考虑基因组中不同位置进化过程的复杂异质性,基因组进化的贝叶斯无限混合模型使强大的系统发育推断成为可能。然而,对于大量的现代数据集,马尔可夫链蒙特卡罗抽样技术的计算负担变得令人望而却步。在这里,我们开发了一个变分贝叶斯程序,以加快广泛使用的PhyloBayes MPI程序,该程序处理氨基酸谱的异质性。该方法不是从后验分布中抽样,而是使用一种称为变分分布的可管理分布来近似(未知)后验分布。通过最小化Kullback-Leibler散度来估计变分分布中的参数。为了检验性能,我们分析了由线粒体、质体编码和核蛋白组成的三个经验数据集。我们的变分方法准确地逼近了对系统发育树、混合物比例和混合物中每个组分的氨基酸倾向的贝叶斯推断,同时使用了较少的计算时间数量级。
The pattern of molecular evolution varies among gene sites and genes in a genome. By taking into account the complex heterogeneity of evolutionary processes among sites in a genome, Bayesian infinite mixture models of genomic evolution enable robust phylogenetic inference. With large modern data sets, however, the computational burden of Markov chain Monte Carlo sampling techniques becomes prohibitive. Here, we have developed a variational Bayesian procedure to speed up the widely used PhyloBayes MPI program, which deals with the heterogeneity of amino acid profiles. Rather than sampling fromthe posterior distribution, the procedure approximates the (unknown) posterior distribution using a manageable distribution called the variational distribution. The parameters in the variational distribution are estimated by minimizing Kullback-Leibler divergence. To examine performance, we analyzed three empirical data sets consisting of mitochondrial, plastid-encoded, and nuclear proteins. Our variational method accurately approximated the Bayesian inference of phylogenetic tree, mixture proportions, and the amino acid propensity of each component of the mixture while using orders of magnitude less computational time.