Computational Performance and Statistical Accuracy of *BEAST and Comparisons with Other Methods.

Computational Performance and Statistical Accuracy of *BEAST and Comparisons with Other Methods.
复制标题

DOI:
10.1093/sysbio/syv118
复制
发表时间:
2016-05
期刊:
影响因子:
6.5
通讯作者:
Drummond AJ
Drummond AJ
中科院分区:
生物学1区
文献类型:
--
作者:
Ogilvie HA;Heled J;Xie D;Drummond AJ

文献摘要

被引文献

相似文献

在多物种结合的分子进化模型下,基因树在一个共享的物种树中具有独立的进化历史。相比之下,超矩阵连接方法假设基因树共享一个共同的系谱历史,从而将基因合并等同于物种分化。以前的研究发现,它的预测分布拟合经验数据的多物种聚结的支持,并串联是不是一个一致的估计的物种树。*BEAST是多物种合并的完全贝叶斯实现,很受欢迎,但计算密集,因此系统发育数据集的大小增加既是计算挑战,也是更好的系统学的机会。使用模拟研究,我们表征了 *BEAST的缩放行为,并能够定量预测增加位点数量对计算性能和统计准确性的影响。在广泛的参数范围内进行的后续模拟表明,随着分支长度的减少和基因座数量的增加,*BEAST相对于串联的统计性能都有所改善。最后,使用基于来自两个基因组数据集的估计参数的模拟,我们比较了一系列物种树和串联方法的性能,以表明使用具有数十个位点的 *BEAST比使用具有数千个位点的串联更可取。我们的研究结果提供了深入了解贝叶斯物种树估计的实用性,所需的位点的数量,以获得一个给定的准确度和超矩阵或总结方法的情况下,将优于完全贝叶斯多物种合并。
Under the multispecies coalescent model of molecular evolution, gene trees have independent evolutionary histories within a shared species tree. In comparison, supermatrix concatenation methods assume that gene trees share a single common genealogical history, thereby equating gene coalescence with species divergence. The multispecies coalescent is supported by previous studies which found that its predicted distributions fit empirical data, and that concatenation is not a consistent estimator of the species tree. *BEAST, a fully Bayesian implementation of the multispecies coalescent, is popular but computationally intensive, so the increasing size of phylogenetic data sets is both a computational challenge and an opportunity for better systematics. Using simulation studies, we characterize the scaling behavior of *BEAST, and enable quantitative prediction of the impact increasing the number of loci has on both computational performance and statistical accuracy. Follow-up simulations over a wide range of parameters show that the statistical performance of *BEAST relative to concatenation improves both as branch length is reduced and as the number of loci is increased. Finally, using simulations based on estimated parameters from two phylogenomic data sets, we compare the performance of a range of species tree and concatenation methods to show that using *BEAST with tens of loci can be preferable to using concatenation with thousands of loci. Our results provide insight into the practicalities of Bayesian species tree estimation, the number of loci required to obtain a given level of accuracy and the situations in which supermatrix or summary methods will be outperformed by the fully Bayesian multispecies coalescent.