Phylogeny Estimation Given Sequence Length Heterogeneity.

Phylogeny Estimation Given Sequence Length Heterogeneity.
复制标题

DOI:
10.1093/sysbio/syaa058
复制
发表时间:
2021-02-10
期刊:
影响因子:
6.5
通讯作者:
Warnow T
Warnow T
中科院分区:
生物学1区
文献类型:
--
作者:
Smirnov V;Warnow T

文献摘要

参考文献

被引文献

相似文献

系统发育评估是许多生物学研究中的重要一步,也面临着许多众所周知的挑战。随着测序技术成本的下降,生物学家现在有越来越多的数据集可用于系统发育评估。在这里,我们解决了在给定具有全长序列和片段序列的组合的大型数据集的情况下估计树的挑战,这可能是由于各种原因而产生的,包括样本收集、测序技术和分析管道。我们比较了两种基本方法:(1)在完整的数据集上计算比对,然后在比对上计算最大似然树,或者(2)在全长序列上构建比对和树,然后使用系统发生位置将剩余的序列(通常是零碎的)添加到树中。我们在一系列模拟数据集上探索了这两种方法,每个模拟数据集有1000个序列,进化速度不同,以及两个生物数据集。我们的研究表明,不同方法之间存在一些显著的性能差异,特别是当存在大量的序列长度异质性和高进化率时。我们特别发现,使用UPP比对序列和使用RAxML计算比对上的树提供了最好的准确性,大大优于使用系统发育放置方法计算的树。我们还发现,FastTree对于包含片段序列的比对的准确性很差。总体而言,我们的研究为比较不同的系统发育估计方法和管道的文献提供了洞察力,并为未来的方法发展提出了方向。[系统发育评估、序列长度异质性、系统发育位置。]
Phylogeny estimation is a major step in many biological studies, and has many well known challenges. With the dropping cost of sequencing technologies, biologists now have increasingly large datasets available for use in phylogeny estimation. Here we address the challenge of estimating a tree given large datasets with a combination of full-length sequences and fragmentary sequences, which can arise due to a variety of reasons, including sample collection, sequencing technologies, and analytical pipelines. We compare two basic approaches: (1) computing an alignment on the full dataset and then computing a maximum likelihood tree on the alignment, or (2) constructing an alignment and tree on the full length sequences and then using phylogenetic placement to add the remaining sequences (which will generally be fragmentary) into the tree. We explore these two approaches on a range of simulated datasets, each with 1000 sequences and varying in rates of evolution, and two biological datasets. Our study shows some striking performance differences between methods, especially when there is substantial sequence length heterogeneity and high rates of evolution. We find in particular that using UPP to align sequences and RAxML to compute a tree on the alignment provides the best accuracy, substantially outperforming trees computed using phylogenetic placement methods. We also find that FastTree has poor accuracy on alignments containing fragmentary sequences. Overall, our study provides insights into the literature comparing different methods and pipelines for phylogenetic estimation, and suggests directions for future method development. [Phylogeny estimation, sequence length heterogeneity, phylogenetic placement.]
EPA-NG:遗传序列的大量平行进化位置。
DOI: 10.1093/sysbio/syy054
发表时间: 2019-03-01
期刊: Systematic biology
影响因子: 6.5
作者:
Barbera P;Kozlov AM;Czech L;Morel B;Darriba D;Flouri T;Stamatakis A
通讯作者: Stamatakis A
DOI: 10.1371/journal.pone.0027731
发表时间: 2011
期刊: PloS one
影响因子: 3.7
作者:
Liu K;Linder CR;Warnow T
通讯作者: Warnow T
DOI: 10.1371/journal.pone.0056859
发表时间: 2013
期刊: PloS one
影响因子: 3.7
作者:
Matsen FA 4th;Evans SN
通讯作者: Evans SN
DOI: 10.1128/msystems.00210-17
发表时间: 2018-01-01
期刊: MSYSTEMS
影响因子: 6.4
作者:
Janssens, Syvie;Schotsaert, Michael;Zwaka, Thomas P.
通讯作者: Zwaka, Thomas P.
DOI: 10.1006/jmbi.1994.1104
发表时间: 1994-02-04
影响因子: 5.6
作者:
KROGH, A;BROWN, M;HAUSSLER, D
通讯作者: HAUSSLER, D