Tail Paradox, Partial Identifiability, and Influential Priors in Bayesian Branch Length Inference

Tail Paradox, Partial Identifiability, and Influential Priors in Bayesian Branch Length Inference
复制标题

DOI:
10.1093/molbev/msr210
复制
发表时间:
2012-01-01
影响因子:
10.7
通讯作者:
Yang, Ziheng
Yang, Ziheng
中科院分区:
生物学1区
文献类型:
--
作者:
Rannala, Bruce;Zhu, Tianqi;Yang, Ziheng

文献摘要

被引文献

相似文献

最近的研究已经观察到,使用MrBayes程序对序列数据集进行贝叶斯分析有时会产生极大的分支长度,树长度(分支长度之和)的后验可信区间不包括最大似然估计。对这一现象的解释包括后验模型中存在多个局部峰值、后验模型尾部的链缺乏收敛、混合问题以及分支长度的先验信息错误。在这里,我们分析了贝叶斯马尔可夫链蒙特卡罗算法的行为时,链是在尾部的后验分布,并注意到,所有这些现象都可能发生。在贝叶斯遗传学中,当分支长度增加到无穷大时,似然函数接近一个常数而不是零。平尾的可能性可能会导致不良的混合和不适当的影响的先验。我们认为,在许多贝叶斯分析中产生的极端分支长度估计的主要原因是当前贝叶斯系统发育程序中分支长度的默认先验选择不佳。MrBayes中的默认先验为分支长度分配独立和相同的分布,对树的长度强加了强(和不合理的)假设。该问题是加剧了强相关性的分支长度和参数之间的可变速率之间的网站或网站分区的模型。为了解决这个问题,我们提出了两个多元先验的分支长度(称为复合狄利克雷先验),是相当分散的,并证明其效用的特殊情况下,分支长度估计的星星的演化。我们的分析突出了需要仔细思考的规范高维先验贝叶斯分析。
Recent studies have observed that Bayesian analyses of sequence data sets using the program MrBayes sometimes generate extremely large branch lengths, with posterior credibility intervals for the tree length (sum of branch lengths) excluding the maximum likelihood estimates. Suggested explanations for this phenomenon include the existence of multiple local peaks in the posterior, lack of convergence of the chain in the tail of the posterior, mixing problems, and misspecified priors on branch lengths. Here, we analyze the behavior of Bayesian Markov chain Monte Carlo algorithms when the chain is in the tail of the posterior distribution and note that all these phenomena can occur. In Bayesian phylogenetics, the likelihood function approaches a constant instead of zero when the branch lengths increase to infinity. The flat tail of the likelihood can cause poor mixing and undue influence of the prior. We suggest that the main cause of the extreme branch length estimates produced in many Bayesian analyses is the poor choice of a default prior on branch lengths in current Bayesian phylogenetic programs. The default prior in MrBayes assigns independent and identical distributions to branch lengths, imposing strong (and unreasonable) assumptions about the tree length. The problem is exacerbated by the strong correlation between the branch lengths and parameters in models of variable rates among sites or among site partitions. To resolve the problem, we suggest two multivariate priors for the branch lengths (called compound Dirichlet priors) that are fairly diffuse and demonstrate their utility in the special case of branch length estimation on a star phylogeny. Our analysis highlights the need for careful thought in the specification of high-dimensional priors in Bayesian analyses.