Information Theoretic Approaches to Whole Genome Phylogenies

Information Theoretic Approaches to Whole Genome Phylogenies
复制标题

全基因组系统发育的信息论方法

DOI:
--
复制
发表时间:
2005
期刊:
Annual International Conference on Research in Computational Molecular Biology
影响因子:
--
通讯作者:
B. Chor
B. Chor
中科院分区:
--
文献类型:
--
作者:
D. Burstein;I. Ulitsky;T. Tuller;B. Chor

文献摘要

被引文献

相似文献

我们描述了一种新的方法,有效地重建系统发育树的基础上,整个基因组或蛋白质组的序列。我们的方法的核心是一个新的衡量序列之间的成对距离,其长度可能会有很大的变化。这种测量是基于信息理论工具(Kullback-Leibler相对熵)。我们提出了一个算法,有效地计算这些距离。该算法使用后缀数组在O(l)时间内计算两个l长序列的距离。它的速度足够快,可以构建数百个物种的基因组树,以及近2000种病毒的基因组森林。初步分析的结果显示出显着的协议与“可接受的系统发育的真相”。为了评估我们的方法,它与一些替代方法一起实施,包括以前在文献中发表的两种方法。比较他们的结果,我们的,使用一个“传统”的树和一个标准的树比较方法,我们的算法改进了“竞争”的大幅利润。
We describe a novel method for efficient reconstruction of phylogenetic trees, based on sequences of whole genomes or proteomes. The core of our method is a new measure of pairwise distances between sequences, whose lengths may greatly vary. This measure is based on information theoretic tools (Kullback-Leibler relative entropy). We present an algorithm for efficiently computing these distances. The algorithm uses suffix arrays to compute the distance of two l long sequences in O(l) time. It is fast enough to enable the construction of the phylogenomic tree for hundreds of species, and the phylogenomic forest for almost two thousand viruses. An initial analysis of the results exhibits a remarkable agreement with “acceptable phylogenetic truth”. To assess our approach, it was implemented together with a number of alternative approaches, including two that were previously published in the literature. Comparing their outcome to ours, using a “traditional” tree and a standard tree comparison method, our algorithm improved upon the “competition” by a substantial margin.