Genome phylogenetic analysis based on extended gene contents

Genome phylogenetic analysis based on extended gene contents
复制标题

DOI:
10.1093/molbev/msh138
复制
发表时间:
2004-07-01
影响因子:
10.7
通讯作者:
Zhang, HM
Zhang, HM
中科院分区:
生物学1区
文献类型:
--
作者:
Gu, X;Zhang, HM

文献摘要

被引文献

相似文献

随着全基因组数据的快速增长,全基因组方法如基因内容成为流行的基因组遗传推断,包括生命树。然而,基因组进化的基本模型是不清楚的,和建议(特设)基因组距离测量可能会违反加和性。在这篇文章中,我们制定了一个随机框架的基因组进化,这提供了一个基础,定义一个加性基因组距离。然而,我们表明,这是很难利用典型的基因含量数据,即,跨基因组的基因家族的存在或不存在来估计基因组距离。我们通过引入扩展基因含量的概念来解决这个问题;即,给定基因组中基因家族的状态可以是缺失、作为单拷贝存在或作为重复存在,其中任何一个都可以用于估计基因组距离和系统发育推断。计算机模拟表明,新的树生成方法是有效的,一致的,相当强大的。35个微生物全基因组的实例表明,它不仅有助于研究普遍的生命树,而且有助于探索基因组的进化模式。
With the rapid growth of entire genome data, whole-genome approaches such as gene content become popular for genome phylogeny inference, including the tree of life. However, the underlying model for genome evolution is unclear, and the proposed (ad hoc) genome distance measure may violate the additivity. In this article, we formulate a stochastic framework for genome evolution, which provides a basis for defining an additive genome distance. However, we show that it is difficult to utilize the typical gene content data-i.e., the presence or absence of gene families across genomes-to estimate the genome distance. We solve this problem by introducing the concept of extended gene content; that is, the status of a gene family in a given genome could be absence, presence as single copy, or presence as duplicates, any of which can be used to estimate the genome distance and phylogenetic inference. Computer simulation shows that the new tree-making method is efficient, consistent, and fairly robust. The example of 35 microbial complete genomes demonstrates that it is useful not only to study the universal tree of life but also to explore the evolutionary pattern of genomes.