Species trees from gene trees: Reconstructing Bayesian posterior distributions of a species phylogeny using estimated gene tree distributions

Species trees from gene trees: Reconstructing Bayesian posterior distributions of a species phylogeny using estimated gene tree distributions
复制标题

DOI:
10.1080/10635150701429982
复制
发表时间:
2007-01-01
期刊:
影响因子:
6.5
通讯作者:
Pearl, Dennis K.
Pearl, Dennis K.
中科院分区:
生物学1区
文献类型:
--
作者:
Liu, Liang;Pearl, Dennis K.

文献摘要

被引文献

相似文献

现在有了相当数量的多位点分子数据,推断一组物种进化历史的愿望应该更加可行。然而,目前的分子系统发育范式仍然重建基因树来表示物种树。此外,通常使用的组合数据的方法,例如级联方法,已知在某些情况下是不一致的。在本文中,我们提出了一个贝叶斯层次模型来估计一组物种的遗传使用多个估计的基因树分布,如在贝叶斯分析的DNA序列数据中出现的。我们的模型采用了传统遗传学中使用的替代模型,但也使用结合理论来解释从物种树到基因树以及从基因树到序列数据的系谱信号,从而形成一个完整的随机模型来同时估计基因树、物种树、祖先群体大小和物种分歧时间。我们的模型是建立在假设基因树,即使是不连锁的位点,是相关的,由于来自一个单一的物种树,因此应该联合估计。我们将该方法应用于两个多位点的DNA序列数据集。物种树拓扑结构和发散时间的估计似乎是强大的人口规模的先验,而有效人口规模的估计是敏感的分析中使用的先验。这些分析还表明,该模型是上级的级联方法,在拟合这些数据集,从而提供了一个更现实的评估,在分布的物种树,可能产生的分子信息在手的变化。我们的模型和算法的未来改进应包括考虑其他因素,可能会导致不一致的基因树和物种树,如水平转移或基因复制。
The desire to infer the evolutionary history of a group of species should be more viable now that a considerable amount of multilocus molecular data is available. However, the current molecular phylogenetic paradigm still reconstructs gene trees to represent the species tree. Further, commonly used methods of combining data, such as the concatenation method, are known to be inconsistent in some circumstances. In this paper, we propose a Bayesian hierarchical model to estimate the phylogeny of a group of species using multiple estimated gene tree distributions, such as those that arise in a Bayesian analysis of DNA sequence data. Our model employs substitution models used in traditional phylogenetics but also uses coalescent theory to explain genealogical signals from species trees to gene trees and from gene trees to sequence data, thereby forming a complete stochastic model to estimate gene trees, species trees, ancestral population sizes, and species divergence times simultaneously. Our model is founded on the assumption that gene trees, even of unlinked loci, are correlated due to being derived from a single species tree and therefore should be estimated jointly. We apply the method to two multilocus data sets of DNA sequences. The estimates of the species tree topology and divergence times appear to be robust to the prior of the population size, whereas the estimates of effective population sizes are sensitive to the prior used in the analysis. These analyses also suggest that the model is superior to the concatenation method in fitting these data sets and thus provides a more realistic assessment of the variability in the distribution of the species tree that may have produced the molecular information at hand. Future improvements of our model and algorithm should include consideration of other factors that can cause discordance of gene trees and species trees, such as horizontal transfer or gene duplication.