Robustness to Divergence Time Underestimation When Inferring Species Trees from Estimated Gene Trees

Robustness to Divergence Time Underestimation When Inferring Species Trees from Estimated Gene Trees
复制标题

DOI:
10.1093/sysbio/syt059
复制
发表时间:
2014-01-01
期刊:
影响因子:
6.5
通讯作者:
Degnan, James H.
Degnan, James H.
中科院分区:
生物学1区
文献类型:
--
作者:
DeGiorgio, Michael;Degnan, James H.

文献摘要

被引文献

相似文献

为了从从多个基因组数据集估计的基因树推断物种树,需要可以处理数十到数百个基因座的易处理的方法。我们研究了几种计算效率高的方法MP-EST,星星,STEAC,STELLS和STEM推断物种树的基因树估计使用最大似然(ML)和贝叶斯方法。在检查的方法中,我们发现基于拓扑的方法通常使用ML基因树表现得更好,使用贝叶斯基因树的方法通常使用合并时间表现得更好,MP-EST,星星,STEAC和STELLS在大多数情况下都优于STEM。我们研究为什么干树(也称为玻璃或最大树)是不太准确的估计基因树比较估计和真实的聚结时间,使用模拟进行物种树推断,并分析一个大猿数据集保持跟踪假阳性和假阴性率推断分支。我们发现,虽然真正的聚结时间比多物种聚结模型下的物种形成时间更古老,估计的聚结时间往往比物种形成时间更近。当用ML估计基因树时,这种低估会导致增加的偏差和缺乏分辨率,增加采样(等位基因或基因座)。使用贝叶斯基因树估计,这个问题似乎不那么严重。
To infer species trees from gene trees estimated from phylogenomic data sets, tractable methods are needed that can handle dozens to hundreds of loci. We examine several computationally efficient approaches-MP-EST, STAR, STEAC, STELLS, and STEM-for inferring species trees from gene trees estimated using maximum likelihood (ML) and Bayesian approaches. Among the methods examined, we found that topology-based methods often performed better using ML gene trees and methods employing coalescent times typically performed better using Bayesian gene trees, with MP-EST, STAR, STEAC, and STELLS outperforming STEM under most conditions. We examine why the STEM tree (also called GLASS or Maximum Tree) is less accurate on estimated gene trees by comparing estimated and true coalescence times, performing species tree inference using simulations, and analyzing a great ape data set keeping track of false positive and false negative rates for inferred clades. We find that although true coalescence times are more ancient than speciation times under the multispecies coalescent model, estimated coalescence times are often more recent than speciation times. This underestimation can lead to increased bias and lack of resolution with increased sampling (either alleles or loci) when gene trees are estimated with ML. The problem appears to be less severe using Bayesian gene-tree estimates.