Species tree inference by minimizing deep coalescences.

Species tree inference by minimizing deep coalescences.
复制标题

DOI:
10.1371/journal.pcbi.1000501
复制
发表时间:
2009-09
影响因子:
4.3
通讯作者:
Nakhleh L
Nakhleh L
中科院分区:
生物学2区
文献类型:
--
作者:
Than C;Nakhleh L

文献摘要

参考文献

被引文献

相似文献

在1997年的一篇开创性论文中,W. Maddison提出了最小化深度合并(MDC),作为从一组不一致的基因树推断物种树的优化标准,假设不一致完全是由于谱系排序。在随后的论文中,Maddison和Knowles提供并实现了一种搜索启发式算法,用于优化MDC标准,给定一组基因树。然而,启发式算法并不能保证计算出最优解,而且它的爬山搜索使它在实践中很慢。在本文中,我们提供了两个精确的解决方案的问题,从一组基因树的MDC标准下推断的物种树。换句话说,我们的解决方案保证从一组基因树中找到最小化深度合并总数的树。一种解决方案是基于一种新的整数线性规划(ILP)的制定,另一种是基于一个简单的动态规划(DP)的方法。强大的ILP求解器,如CPLEX,使第一个解决方案有吸引力,特别是对于非常大规模的问题,而基于DP的解决方案消除了对专有工具的依赖,其简单性使得它很容易与其他可能导致基因树不一致的基因组事件集成。使用精确的解决方案,我们分析了106个位点的数据集,从8个酵母物种,268个位点的数据集,从8个顶复门物种,和几个模拟数据集。我们表明,MDC标准提供了非常准确的估计的物种树拓扑结构,我们的解决方案是非常快的,从而允许基因组规模的数据集的准确分析。此外,解决方案的效率允许快速探索次优解决方案,这对于基于简约的标准(如MDC)非常重要,正如我们所示。我们发现,搜索的物种树的兼容性图中的基因树引起的集群可能是足够的,在实践中,这一发现有助于改善优化解决方案的计算要求。进一步,我们研究了MDC准则的统计一致性和收敛速度,以及它在物种树推断中的最优性。最后,我们展示了我们的解决方案如何用于识别可能导致数据中某些不一致的潜在水平基因转移事件,从而增强Maddison的原始框架。我们在PhyloNet软件包中实施了我们的解决方案,该软件包可在www.example.com免费获得。推断一组物种的进化历史,被称为物种树,是生物学和其他领域中最重要的任务。从分子序列中完成这项任务的传统方法需要对所考虑的物种中的一个基因进行测序,重建基因的进化历史,并将其宣布为物种树。然而,由于测序技术的进步,最近对多个基因数据集的分析表明,同一组物种的基因树可能彼此不一致,也可能与物种树不一致。因此,尽管有这样的分歧,推断的种树的方法的发展是势在必行。在本文中,我们提出了这样一种方法,它寻求的树,最大限度地减少输入集的基因树和推断之间的分歧。我们已经实现了我们的方法,并研究了它的性能,在精度和计算效率方面,两个生物数据集和大量的模拟数据集。我们的分析,生物和合成数据集,表明高精度的方法,以及在实践中的计算效率的解决方案。因此,我们的方法是一个很好的候选人推断准确的物种树,尽管基因树的分歧,在基因组规模。
In a 1997 seminal paper, W. Maddison proposed minimizing deep coalescences, or MDC, as an optimization criterion for inferring the species tree from a set of incongruent gene trees, assuming the incongruence is exclusively due to lineage sorting. In a subsequent paper, Maddison and Knowles provided and implemented a search heuristic for optimizing the MDC criterion, given a set of gene trees. However, the heuristic is not guaranteed to compute optimal solutions, and its hill-climbing search makes it slow in practice. In this paper, we provide two exact solutions to the problem of inferring the species tree from a set of gene trees under the MDC criterion. In other words, our solutions are guaranteed to find the tree that minimizes the total number of deep coalescences from a set of gene trees. One solution is based on a novel integer linear programming (ILP) formulation, and another is based on a simple dynamic programming (DP) approach. Powerful ILP solvers, such as CPLEX, make the first solution appealing, particularly for very large-scale instances of the problem, whereas the DP-based solution eliminates dependence on proprietary tools, and its simplicity makes it easy to integrate with other genomic events that may cause gene tree incongruence. Using the exact solutions, we analyze a data set of 106 loci from eight yeast species, a data set of 268 loci from eight Apicomplexan species, and several simulated data sets. We show that the MDC criterion provides very accurate estimates of the species tree topologies, and that our solutions are very fast, thus allowing for the accurate analysis of genome-scale data sets. Further, the efficiency of the solutions allow for quick exploration of sub-optimal solutions, which is important for a parsimony-based criterion such as MDC, as we show. We show that searching for the species tree in the compatibility graph of the clusters induced by the gene trees may be sufficient in practice, a finding that helps ameliorate the computational requirements of optimization solutions. Further, we study the statistical consistency and convergence rate of the MDC criterion, as well as its optimality in inferring the species tree. Finally, we show how our solutions can be used to identify potential horizontal gene transfer events that may have caused some of the incongruence in the data, thus augmenting Maddison's original framework. We have implemented our solutions in the PhyloNet software package, which is freely available at: http://bioinfo.cs.rice.edu/phylonet. Inferring the evolutionary history of a set of species, known as the species tree, is a task of utmost significance in biology and beyond. The traditional approach to accomplishing this task from molecular sequences entails sequencing a gene in the set of species under consideration, reconstructing the gene's evolutionary history, and declaring it to be the species tree. However, recent analyses of multiple gene data sets, made available thanks to advances in sequencing technologies, have indicated that gene trees in the same group of species may disagree with each other, as well as with the species tree. Therefore, the development of methods for inferring the species tree despite such disagreements is imperative. In this paper, we propose such a method, which seeks the tree that minimizes the amount of disagreement between the input set of gene trees and the inferred one. We have implemented our method and studied its performance, in terms of accuracy and computational efficiency, on two biological data sets and a large number of simulated data sets. Our analyses, of both the biological and synthetic data sets, indicate high accuracy of the method, as well as computationally efficient solutions in practice. Hence, our method makes a good candidate for inferring accurate species trees, despite gene tree disagreements, at a genomic scale.
DOI: 10.1093/molbev/msn213
发表时间: 2008-12
影响因子: 10.7
作者:
Kuo, Chih-Horng;Wares, John P.;Kissinger, Jessica C.
通讯作者: Kissinger, Jessica C.
DOI: 10.1371/journal.pgen.0020173
发表时间: 2006-10-27
期刊: PLoS genetics
影响因子: 4.5
作者:
Pollard DA;Iyer VN;Moses AM;Eisen MB
通讯作者: Eisen MB
DOI: 10.1073/pnas.0607004104
发表时间: 2007-04-03
影响因子: 11.1
作者:
Edwards, Scott V.;Liu, Liang;Pearl, Dennis K.
通讯作者: Pearl, Dennis K.
DOI: 10.1093/sysbio/46.3.523
发表时间: 1997-09-01
期刊: SYSTEMATIC BIOLOGY
影响因子: 6.5
作者:
Maddison, WP
通讯作者: Maddison, WP
DOI: 10.1080/10635150701429982
发表时间: 2007-01-01
期刊: SYSTEMATIC BIOLOGY
影响因子: 6.5
作者:
Liu, Liang;Pearl, Dennis K.
通讯作者: Pearl, Dennis K.