Modeling compositional heterogeneity

Modeling compositional heterogeneity
复制标题

DOI:
10.1080/10635150490445779
复制
发表时间:
2004-06-01
期刊:
影响因子:
6.5
通讯作者:
Foster, PG
Foster, PG
中科院分区:
生物学1区
文献类型:
--
作者:
Foster, PG

文献摘要

被引文献

相似文献

谱系之间的成分异质性可能会损害系统发育分析,因为常用的模型假设成分同质的数据。模型,可以容纳成分的异质性与一些额外的参数在这里描述,并在两个例子中,真正的树是已知的信心。它示出使用似然比检验,可以实现充分的成分异质性的建模与几个组成参数,数据可能不需要与单独的组成参数为树中的每个分支建模。树搜索和放置在树上的组成向量的贝叶斯框架中使用马尔可夫链蒙特卡罗(MCMC)方法。在最大似然(ML)和贝叶斯框架中对模型与数据的拟合进行评估。在ML框架中,使用Goldman-Cox检验评估整体模型拟合,并使用新的基于树和模型的组合拟合检验评估(可能是异构的)模型所暗示的组合与数据组合的拟合。在贝叶斯框架中,使用后验预测模拟评估整体模型拟合和组成拟合。结果表明,当组成不被容纳,然后该模型不适合,并发现不正确的树,但当组成被容纳,然后该模型适合,并获得已知的正确的树。
Compositional heterogeneity among lineages can compromise phylogenetic analyses, because models in common use assume compositionally homogeneous data. Models that can accommodate compositional heterogeneity with few extra parameters are described here, and used in two examples where the true tree is known with confidence. It is shown using likelihood ratio tests that adequate modeling of compositional heterogeneity can be achieved with few composition parameters, that the data may not need to be modelled with separate composition parameters for each branch in the tree. Tree searching and placement of composition vectors on the tree are done in a Bayesian framework using Markov chain Monte Carlo (MCMC) methods. Assessment of fit of the model to the data is made in both maximum likelihood (ML) and Bayesian frameworks. In an ML framework, overall model fit is assessed using the Goldman-Cox test, and the fit of the composition implied by a (possibly heterogeneous) model to the composition of the data is assessed using a novel tree- and model-based composition fit test. In a Bayesian framework, overall model fit and composition fit are assessed using posterior predictive simulation. It is shown that when composition is not accommodated, then the model does not fit, and incorrect trees are found; but when composition is accommodated, the model then fits, and the known correct phylogenies are obtained.