Tutorial on Phylogenetic Tree Estimation

Tutorial on Phylogenetic Tree Estimation
复制标题

系统发育树估计教程

DOI:
--
复制
发表时间:
1999
期刊:
Intelligent Systems in Molecular Biology
影响因子:
--
通讯作者:
Junhyong Kim
Junhyong Kim
中科院分区:
--
文献类型:
--
作者:
Junhyong Kim

文献摘要

参考文献

被引文献

相似文献

所有的生物学学科都被物种有着共同的历史这一观点所统一。生命的谱系历史也被称为“进化树”通常用一棵分叉的、标有叶子的树来表示。进化树的使用是许多生物学问题的基础步骤,例如多序列比对、蛋白质结构和功能预测以及药物设计。系统发育研究的主要科学目标不是解决给定的优化问题,而是恢复由真实进化树的拓扑结构所代表的物种形成或基因复制事件的顺序。(定位进化树的根是一项科学上的艰巨任务,因此,如果一种方法恢复了无根树的拓扑结构,则该方法被认为是成功的。这意味着关于优化问题的好的或差的性能仅在它保证关于拓扑估计的好的或差的性能的程度上是重要的。不幸的是,由于几个原因,推断进化树是一个非常困难的问题。首先,遗传问题是一个复杂的统计问题,因为它的参数空间具有复杂的结构,并且没有任何货架解决方案可以应用于遗传问题。系统发育问题也提出了相当大的计算挑战。典型的数据集现在包括几百个物种,目前可用的树重建方法是不足以分析这样的数据集的任务。例如,500株植物的rbcL DNA序列数据集已经分析了好几年,但没有解决方案。为什么这些分析是如此复杂的解释很简单:优化问题是NP难的,并且试图解决这些优化问题所使用的算法使用爬山技术来搜索系统发育树的指数大空间。系统发育重建的统计方法对进化过程进行了随机建模,并研究了在不同模型树下生成的nite长度序列数据集上恢复系统发育树的方法的准确性。这些研究表明,一旦序列足够长,一些方法以高概率恢复真正的树拓扑,而其他方法则没有这样的保证。在过去的十年左右,计算机科学家也开始设计和分析系统发育方法在这些统计模型下的性能。这种兴趣的结果之一是使用统计模型的.
1 Tutorial Summary All biological disciplines are united by the idea that species share a common history. The genealogical history of life-also called an \evolutionary tree"-is usually represented by a bifurcating, leaf-labeled tree. The use of evolutionary trees is a fundamental step in many biological problems, such as multiple sequence alignments, protein structure and function prediction, and drug design. The primary scientiic objective of phylogenetic studies is not to solve a given optimization problem, but rather to recover the order of speciation or gene duplication events represented by the topology of the true evolutionary tree. (Locating the root of the evolutionary tree is a scientiically diicult task, so that a method is considered to have been successful if it recovers the topology of the unrooted tree.) This means that good or poor performance with respect to optimization problems is only important to the degree that it guarantees good or poor performance with respect to topology estimation. Unfortunately, inferring evolutionary trees is an enormously diicult problem for several reasons. For one, the phylogeny problem is a diicult statistical problem because its parameter space has a complicated structure, and there is nòoo the shelf' solution to the phylogeny problem that can be applied. The phylogeny problem also presents a considerable computational challenge. Typical data sets now consist of several hundred species, and presently available tree reconstruction methods are inadequate to the task of analyzing such datasets. For example, an rbcL DNA sequence data set of 500 plants has been analyzed for several years now, without solution. The explanation for why these analyses are so diicult is simple: the optimization problems are NP-hard, and the heuristics used in an attempt to solve these optimization problems use hill-climbing techniques to search through an exponentially large space of phylogenetic trees. Statistical approaches towards phylogeny reconstruction have modeled the evolutionary process stochasti-cally, and have studied the performance of methods for recovering phylogenetic trees in terms of the accuracy of these methods on datasets of nite length sequences generated under diierent model trees. These studies have shown that some methods recover the true tree topology with high probability, once the sequences are long enough, while other methods have no such guarantees. Over the last decade or so, computer scientists have also begun to design and analyze the performance of phylogenetic methods under these statistical models. One of the results of this interest in using statistical models of …
DOI: 10.1073/pnas.94.13.6585
发表时间: 1997
影响因子: 11.1
作者:
Warnow,T
通讯作者: Warnow,T
DOI: 10.1006/jmbi.1994.1104
发表时间: 1994-02-04
影响因子: 5.6
作者:
KROGH, A;BROWN, M;HAUSSLER, D
通讯作者: HAUSSLER, D
DOI: 10.1126/science.1590849
发表时间: 1992-02-07
期刊: SCIENCE
影响因子: 56.9
作者:
TEMPLETON, AR
通讯作者: TEMPLETON, AR
DOI: 10.1093/oxfordjournals.molbev.a025575
发表时间: 1996-01-01
影响因子: 10.7
作者:
Felsenstein, J;Churchill, GA
通讯作者: Churchill, GA