Naive binning improves phylogenomic analyses

Naive binning improves phylogenomic analyses
复制标题

DOI:
10.1093/bioinformatics/btt394
复制
发表时间:
2013-09-15
期刊:
影响因子:
5.8
通讯作者:
Warnow, Tandy
Warnow, Tandy
中科院分区:
生物学3区
文献类型:
--
作者:
Bayzid, Md Shamsuzzoha;Warnow, Tandy

文献摘要

被引文献

相似文献

动机:在不完全谱系分类(ILS)的存在下的物种树估计是一个主要的挑战,为植物基因组分析。虽然已经开发了许多方法,这个问题,很少了解这些方法的相对性能估计基因树时估计不佳,由于系统发育signal.Results不足:我们探讨了一些方法的性能估计物种树从多个标记的模拟数据集,基因树不同的物种树,由于ILS。我们包括 *BEAST,级联分析和几个“总结方法”:BUCKy,MP-EST,最小化深度合并,简约矩阵表示和贪婪共识。我们发现,*BEAST和拼接给出了非常好的结果,通常比其他方法具有更高的准确性。我们观察到,*BEAST的准确性很大程度上是由于它能够共同估计基因树和物种树。然而,*BEAST是计算密集型的,这使得在具有100个或更多基因或20个以上分类群的数据集上运行具有挑战性。我们提出了一种新的方法,物种树估计中的基因被划分成集,和物种树估计所产生的“超基因”。我们表明,这种技术提高了 *BEAST的可扩展性,而不影响其准确性,并提高了摘要方法的准确性。因此,在ILS存在的情况下,初始分箱可以改善PCR基因组分析。
Motivation: Species tree estimation in the presence of incomplete lineage sorting (ILS) is a major challenge for phylogenomic analysis. Although many methods have been developed for this problem, little is understood about the relative performance of these methods when estimated gene trees are poorly estimated, owing to inadequate phylogenetic signal.Results: We explored the performance of some methods for estimating species trees from multiple markers on simulated datasets in which gene trees differed from the species tree owing to ILS. We included *BEAST, concatenated analysis and several 'summary methods': BUCKy, MP-EST, minimize deep coalescence, matrix representation with parsimony and the greedy consensus. We found that *BEAST and concatenation gave excellent results, often with substantially improved accuracy over the other methods. We observed that *BEAST's accuracy is largely due to its ability to co-estimate the gene trees and species tree. However, *BEAST is computationally intensive, making it challenging to run on datasets with 100 or more genes or with more than 20 taxa. We propose a new approach to species tree estimation in which the genes are partitioned into sets, and the species tree is estimated from the resultant 'supergenes'. We show that this technique improves the scalability of *BEAST without affecting its accuracy and improves the accuracy of the summary methods. Thus, naive binning can improve phylogenomic analysis in the presence of ILS.