APPLES: Fast Distance-based Phylogenetic Placement

APPLES: Fast Distance-based Phylogenetic Placement
复制标题

DOI:
10.1101/475566
复制
发表时间:
2018-11
期刊:
bioRxiv
影响因子:
--
通讯作者:
M. Balaban;Shahab Sarmashghi;S. Mirarab
M. Balaban;Shahab Sarmashghi;S. Mirarab
中科院分区:
其他
文献类型:
--
作者:
M. Balaban;Shahab Sarmashghi;S. Mirarab

文献摘要

相似文献

系统发育放置包括将查询物种添加到现有的系统发育中,并且随着序列数据集的大小和多样性不断增长,其相关性越来越大。放置对于更新现有系统发育以及使用(元)条形码或宏基因组学进行分类学识别样本非常有用。存在系统发育放置的最大似然 (ML) 方法,但这些方法无法扩展到具有数千个叶子的树。它们还依赖于参考树和查询的组装和比对序列,因此无法分析最近在基因组撇读等应用中使用的未组装读数。在这里,我们介绍 APPLES,一种基于距离的系统发育放置方法,它在速度和内存方面将 ML 提高了一个数量级以上,并且在准确性方面非常接近 ML。 APPLES 在数千种树木上的放置比 ML 具有更好的准确性,并且可以在 ML 无法运行的十万种树木上放置。最后,APPLES 可以使用基于 k 聚体的距离准确识别没有组装序列作为参考或查询的样本,这是 ML 无法处理的场景。
Phylogenetic placement consists of adding a query species onto an existing phylogeny and has increasing relevance as sequence datasets continue to grow in size and diversity. Placement is useful for updating existing phylogenies and for identifying samples taxonomically using (meta-)barcoding or metagenomics. Maximum likelihood (ML) methods of phylogenetic placement exist, but these methods are not scalable to trees with many thousands of leaves. They also rely on assembled and aligned sequences for the reference tree and the query and thus cannot analyze unassembled reads used recently in applications such as genome skimming. Here, we introduce APPLES, a distance-based method of phylogenetic placement that improves on ML by more than an order of magnitude in speed and memory and comes very close to ML in accuracy. APPLES has better accuracy than ML for placing on trees with thousands of species and can place on trees with a hundred thousands species where ML cannot run. Finally, APPLES can accurately identify samples without assembled sequences for the reference or the query using k-mer-based distances, a scenario that ML cannot handle.