Choosing among partition models in Bayesian phylogenetics.

Choosing among partition models in Bayesian phylogenetics.
复制标题

DOI:
10.1093/molbev/msq224
复制
发表时间:
2011-01
影响因子:
10.7
通讯作者:
Lewis PO
Lewis PO
中科院分区:
生物学1区
文献类型:
--
作者:
Fan Y;Wu R;Chen MH;Kuo L;Lewis PO

文献摘要

参考文献

被引文献

相似文献

贝叶斯系统发育分析通常依赖于贝叶斯因子(BF)来确定划分数据的最佳方式。反过来,用于计算BF的边际概率通常是使用调和平均(HM)方法估计的,这种方法已被证明是不准确的。我们描述了一种新的更精确的估计模型边际似然的方法,并在模拟数据和经验数据上与HM方法进行了比较。新方法推广了我们前面描述的踏脚石(SS)方法,它利用了使用来自后验分布的样本进行参数化的参考分布。这避免了原始SS方法的一个具有挑战性的方面,即需要从(在Kullback-Leibler意义上)接近先验的分布进行抽样。我们具体讨论了分区模型的选择,并发现使用HM方法可能会导致对过度分区模型的强烈偏好。与HM方法和原始SS方法相比,我们用模拟数据证明了广义SS方法比HM方法有更高的精度(相同数据和划分模型的可重复BF值),并得到比HM方法更合理的BF值。HM方法和广义SS方法在经验数据集上的比较表明,广义SS方法倾向于选择更简单的基于分子进化模式的更符合预期的划分方案。广义SS方法与热力学积分一样,除了后验分布外,还需要从一系列分布中抽样。这种专门的基于路径的马尔可夫链蒙特卡罗分析似乎是准确估计边际可能性的成本。
Bayesian phylogenetic analyses often depend on Bayes factors (BFs) to determine the optimal way to partition the data. The marginal likelihoods used to compute BFs, in turn, are most commonly estimated using the harmonic mean (HM) method, which has been shown to be inaccurate. We describe a new more accurate method for estimating the marginal likelihood of a model and compare it with the HM method on both simulated and empirical data. The new method generalizes our previously described stepping-stone (SS) approach by making use of a reference distribution parameterized using samples from the posterior distribution. This avoids one challenging aspect of the original SS method, namely the need to sample from distributions that are close (in the Kullback–Leibler sense) to the prior. We specifically address the choice of partition models and find that using the HM method can lead to a strong preference for an overpartitioned model. In contrast to the HM method and the original SS method, we show using simulated data that the generalized SS method is strikingly more precise (repeatable BF values of the same data and partition model) and yields BF values that are much more reasonable than those produced by the HM method. Comparisons of HM and generalized SS methods on an empirical data set demonstrate that the generalized SS method tends to choose simpler partition schemes that are more in line with expectation based on inferred patterns of molecular evolution. The generalized SS method shares with thermodynamic integration the need to sample from a series of distributions in addition to the posterior. Such dedicated path-based Markov chain Monte Carlo analyses appear to be a cost of estimating marginal likelihoods accurately.
DOI: 10.1214/aos/1056562461
发表时间: 2003-06-01
影响因子: 4.5
作者:
Neal, RM
通讯作者: Neal, RM
DOI: 10.1109/tac.1974.1100705
发表时间: 1974-01-01
影响因子: 6.8
作者:
AKAIKE, H
通讯作者: AKAIKE, H
DOI: 10.1080/10635150701546249
发表时间: 2007-01-01
期刊: SYSTEMATIC BIOLOGY
影响因子: 6.5
作者:
Brown, Jeremy M.;Lemmon, Alan R.
通讯作者: Lemmon, Alan R.
DOI: 10.1093/molbev/msh123
发表时间: 2004-06-01
影响因子: 10.7
作者:
Huelsenbeck, JP;Larget, B;Alfaro, ME
通讯作者: Alfaro, ME
DOI: 10.2307/2291073
发表时间: 1995-06-01
影响因子: 3.7
作者:
VERDINELLI, I;WASSERMAN, L
通讯作者: WASSERMAN, L