SDM: A fast distance-based approach for (super) tree building in phylogenomics

SDM: A fast distance-based approach for (super) tree building in phylogenomics
复制标题

DOI:
10.1080/10635150600969872
复制
发表时间:
2006-01-01
期刊:
影响因子:
6.5
通讯作者:
Gascuel, Olivier
Gascuel, Olivier
中科院分区:
生物学1区
文献类型:
--
作者:
Criscuolo, Alexis;Berry, Vincent;Gascuel, Olivier

文献摘要

被引文献

相似文献

系统基因组学研究的目的是从大量的同源基因中建立同源基因。这种“基因组大小”的数据需要快速的方法,因为通常大量的类群检查。在这个框架中,基于距离的方法是有用的探索性研究和建立一个起始树,以细化一个更强大的最大似然(ML)的方法。然而,估计进化距离直接从级联基因给穷人的拓扑信号,基因以不同的速度进化。我们提出了一种新的方法,命名为超距离矩阵(SDM),它遵循相同的线平均共识超树(ACS; Lapointe和Cucumel,1997年),并结合从每个基因获得的进化距离到一个单一的距离超矩阵进行分析,使用一个标准的基于距离的算法。SDM在不修改其拓扑信息的情况下使源矩阵变形,以使它们尽可能地彼此接近;然后对这些变形的矩阵进行平均以获得距离超矩阵。我们表明,这个问题是等价的最小二乘准则的线性约束下的最小化。这个问题有一个唯一的解决方案,这是通过解决一个线性系统。由于这个系统是稀疏的,它的实际解决需要O(n(a)k(a))时间,其中n是分类群的数量,k是矩阵的数量,并且a < 2,这使得距离超矩阵可以快速获得。SDM的几种用途,提出了从快速探索性的研究,更准确的方法,需要更重的计算时间。使用模拟,我们表明,SDM是一个相关的替代标准的矩阵表示与简约(MRP)方法,特别是当不同基因的类群集有低重叠。我们还表明,SDM可以用来建立一个很好的开始树的ML方法,这既减少了计算时间,提高了拓扑精度。我们使用SDM来分析Gatesy等人(2002,Syst.Biol.51:652-664)的数据集,该数据集涉及75种胎盘哺乳动物的48个基因。结果表明,这些基因具有较强的速率异质性,证实了模拟的结论。
Phylogenomic studies aim to build phylogenies from large sets of homologous genes. Such "genome-sized" data require fast methods, because of the typically large numbers of taxa examined. In this framework, distance-based methods are useful for exploratory studies and building a starting tree to be refined by a more powerful maximum likelihood (ML) approach. However, estimating evolutionary distances directly from concatenated genes gives poor topological signal as genes evolve at different rates. We propose a novel method, named super distance matrix (SDM), which follows the same line as average consensus supertree (ACS; Lapointe and Cucumel, 1997) and combines the evolutionary distances obtained from each gene into a single distance supermatrix to be analyzed using a standard distance-based algorithm. SDM deforms the source matrices, without modifying their topological message, to bring them as close as possible to each other; these deformed matrices are then averaged to obtain the distance supermatrix. We show that this problem is equivalent to the minimization of a least-squares criterion subject to linear constraints. This problem has a unique solution which is obtained by resolving a linear system. As this system is sparse, its practical resolution requires O(n(a)k(a)) time, where n is the number of taxa, k the number of matrices, and a < 2, which allows the distance supermatrix to be quickly obtained. Several uses of SDM are proposed, from fast exploratory studies to more accurate approaches requiring heavier computing time. Using simulations, we show that SDM is a relevant alternative to the standard matrix representation with parsimony (MRP) method, notably when the taxa sets of the different genes have low overlap. We also show that SDM can be used to build an excellent starting tree for an ML approach, which both reduces the computing time and increases the topogical accuracy. We use SDM to analyze the data set of Gatesy et al. ( 2002, Syst. Biol. 51: 652-664) that involves 48 genes of 75 placental mammals. The results indicate that these genes have strong rate heterogeneity and confirm the simulation conclusions.