Distance-Based Phylogenetic Placement with Statistical Support.

Distance-Based Phylogenetic Placement with Statistical Support.
复制标题

DOI:
10.3390/biology11081212
复制
发表时间:
2022-08-12
期刊:
影响因子:
4.2
通讯作者:
--
中科院分区:
生物学3区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

系统发育定位旨在为现有骨干树上的一个新的查询物种找到最佳位置。快速且准确的基于距离的系统发育定位方法缺乏估计查询序列各种定位的支持值这一关键特征。本研究提出了用于测量基于距离的系统发育定位的支持值的参数方法和非参数方法。 在现代生态学研究中,通常会尝试通过将未知序列置于树上进行系统发育鉴定。此类定位通常从不完整且有噪声的数据中获得,因此有必要用某种不确定性的概念来增强结果。虽然为定位而设计的标准基于似然的方法自然地提供了这种不确定性的度量,但更新且更具扩展性的基于距离的方法缺乏这一关键特征。在此,我们采用了几种参数和非参数抽样方法来测量使用距离获得的系统发育定位的支持度。通过比较不同策略,我们得出结论:非参数自助法比其他方法更准确。我们接着展示了如何使用线性代数公式高效地进行自助法,使其速度提高多达30倍,并将这个优化版本作为基于距离的定位软件APPLES的一部分来实现。通过研究广泛的应用,我们表明最大似然(ML)支持值相对于基于距离的方法的相对准确性取决于应用和数据集。ML对于片段式查询是有利的,而基于距离的支持值对于全长和多基因数据集更准确。通过对不确定性的量化,我们的工作填补了一个关键空白,该空白阻碍了基于距离的定位工具的更广泛应用。
Phylogenetic placement seeks to find the optimal position for a new query species on an existing backbone tree. Fast and accurate distance-based phylogenetic placement methods lack the crucial feature of estimating the support values for various placements of a query sequence. This study presents both parametric and nonparametric methods for measuring the support values of distance-based phylogenetic placements. Phylogenetic identification of unknown sequences by placing them on a tree is routinely attempted in modern ecological studies. Such placements are often obtained from incomplete and noisy data, making it essential to augment the results with some notion of uncertainty. While the standard likelihood-based methods designed for placement naturally provide such measures of uncertainty, the newer and more scalable distance-based methods lack this crucial feature. Here, we adopt several parametric and nonparametric sampling methods for measuring the support of phylogenetic placements that have been obtained with the use of distances. Comparing the alternative strategies, we conclude that nonparametric bootstrapping is more accurate than the alternatives. We go on to show how bootstrapping can be performed efficiently using a linear algebraic formulation that makes it up to 30 times faster and implement this optimized version as part of the distance-based placement software APPLES. By examining a wide range of applications, we show that the relative accuracy of maximum likelihood (ML) support values as compared to distance-based methods depends on the application and the dataset. ML is advantageous for fragmentary queries, while distance-based support values are more accurate for full-length and multi-gene datasets. With the quantification of uncertainty, our work fills a crucial gap that prevents the broader adoption of distance-based placement tools.
DOI: 10.1073/pnas.89.22.10915
发表时间: 1992-11-15
影响因子: 11.1
作者:
HENIKOFF, S;HENIKOFF, JG
通讯作者: HENIKOFF, JG
EPA-NG:遗传序列的大量平行进化位置。
DOI: 10.1093/sysbio/syy054
发表时间: 2019-03-01
期刊: Systematic biology
影响因子: 6.5
作者:
Barbera P;Kozlov AM;Czech L;Morel B;Darriba D;Flouri T;Stamatakis A
通讯作者: Stamatakis A
DOI: 10.1089/106652702761034136
发表时间: 2002-01-01
影响因子: 1.7
作者:
Desper, R;Gascuel, O
通讯作者: Gascuel, O
DOI: 10.1126/science.155.3760.279
发表时间: 1967-01-01
期刊: SCIENCE
影响因子: 56.9
作者:
FITCH, WM;MARGOLIASH, E
通讯作者: MARGOLIASH, E
DOI: 10.1080/10635150600755453
发表时间: 2006-08-01
期刊: SYSTEMATIC BIOLOGY
影响因子: 6.5
作者:
Anisimova, Maria;Gascuel, Olivier
通讯作者: Gascuel, Olivier