Efficient comparative phylogenetics on large trees

Efficient comparative phylogenetics on large trees
复制标题

DOI:
10.1093/bioinformatics/btx701
复制
发表时间:
2018-03-15
期刊:
影响因子:
5.8
通讯作者:
Doebeli, Michael
Doebeli, Michael
中科院分区:
生物学3区
文献类型:
--
作者:
Louca, Stilianos;Doebeli, Michael

文献摘要

被引文献

相似文献

动机:生物多样性数据库现在包含数十万个序列和性状记录。例如,开放生命树包括超过 1,491,000 个后生动物和超过 300,000 个细菌类群。这些数据为分析系统发育性状分布和重建祖先生物多样性提供了独特的机会。然而,现有的比较系统发育工具对于如此大的树木来说扩展性很差,几乎无法使用。结果:在这里,我们提出了一个新的 R 包,名为“castor”,用于对包含数百万个尖端的大树进行比较系统发育。在大树上,蓖麻的速度通常比现有工具快 100-1000 倍。
Motivation: Biodiversity databases now comprise hundreds of thousands of sequences and trait records. For example, the Open Tree of Life includes over 1 491 000 metazoan and over 300 000 bacterial taxa. These data provide unique opportunities for analysis of phylogenetic trait distribution and reconstruction of ancestral biodiversity. However, existing tools for comparative phylogenetics scale poorly to such large trees, to the point of being almost unusable.Results: Here we present a new R package, named 'castor', for comparative phylogenetics on large trees comprising millions of tips. On large trees castor is often 100-1000 times faster than existing tools.