Phylogeny Reconstruction with Alignment-Free Method That Corrects for Horizontal Gene Transfer.

Phylogeny Reconstruction with Alignment-Free Method That Corrects for Horizontal Gene Transfer.
复制标题

DOI:
10.1371/journal.pcbi.1004985
复制
发表时间:
2016-06
影响因子:
4.3
通讯作者:
Otwinowski Z
Otwinowski Z
中科院分区:
生物学2区
文献类型:
--
作者:
Bromberg R;Grishin NV;Otwinowski Z

文献摘要

被引文献

相似文献

测序技术的进步产生了大量的完整基因组。传统上,系统发育分析依赖于同源物的比对,但定义同源物并将它们从同源物中分离出来是一项复杂的任务,可能并不总是适合未来的大型数据集。替代传统的、基于比对的方法是全基因组、无比对方法。这些方法是可扩展的,并且需要最少的人工干预。我们开发了SlopeTree,这是一种新的无对齐方法,通过测量精确子串匹配的衰减作为匹配长度的函数来估计进化距离。SlopeTree可以校正水平基因转移、成分变异和低复杂度序列,以及由同一位点的多个突变引起的分支长度非线性。我们对495种细菌、73种古细菌和72种大肠杆菌和志贺氏菌进行了测试。我们将我们的树与NCBI分类法、基于串联比对的树以及由其他不需要比对的方法生成的树进行了比较。结果与目前关于原核生物进化的知识一致。我们评估了不同方法和设置下树拓扑结构的差异,发现大多数细菌和古细菌都有一组核心蛋白质,这些蛋白质是通过遗传进化而来的。在由完整的基因组而不是核心基因组成的树中,我们观察到一些按表型而不是按系统发育进行分组,例如,一群嗜硫的嗜热细菌聚集在一起,而不考虑它们的门。SlopeTree的源代码可在:http://prodata.swmed.edu/download/pub/slopetree_v1/slopetree.tar.gz。由于缺乏明显的形态特征,细菌和古细菌极难分类,直到技术发展到获得它们的DNA序列;然后可以比较这些序列来估计进化关系。现在,由于技术的进步,从各种各样的生物体中有大量可用的序列。这些进步刺激了算法的发展,这些算法可以使用整个基因组来估计进化关系,而不是更传统的方法,早期使用单个基因,现在通常使用保守基因群。然而,在试图推断进化关系时,存在许多挑战,特别是水平基因转移,即DNA从一个生物体转移到另一个生物体,导致生物体基因组中包含的DNA不能反映其血统进化。我们开发了一种新的全基因组方法来估计进化距离,以识别和纠正水平转移。我们发现,对于我们应用的SlopeTree和所有其他全基因组方法,水平转移导致一些进化距离被严重低估,我们的修正纠正了这一点。
Advances in sequencing have generated a large number of complete genomes. Traditionally, phylogenetic analysis relies on alignments of orthologs, but defining orthologs and separating them from paralogs is a complex task that may not always be suited to the large datasets of the future. An alternative to traditional, alignment-based approaches are whole-genome, alignment-free methods. These methods are scalable and require minimal manual intervention. We developed SlopeTree, a new alignment-free method that estimates evolutionary distances by measuring the decay of exact substring matches as a function of match length. SlopeTree corrects for horizontal gene transfer, for composition variation and low complexity sequences, and for branch-length nonlinearity caused by multiple mutations at the same site. We tested SlopeTree on 495 bacteria, 73 archaea, and 72 strains of Escherichia coli and Shigella. We compared our trees to the NCBI taxonomy, to trees based on concatenated alignments, and to trees produced by other alignment-free methods. The results were consistent with current knowledge about prokaryotic evolution. We assessed differences in tree topology over different methods and settings and found that the majority of bacteria and archaea have a core set of proteins that evolves by descent. In trees built from complete genomes rather than sets of core genes, we observed some grouping by phenotype rather than phylogeny, for instance with a cluster of sulfur-reducing thermophilic bacteria coming together irrespective of their phyla. The source-code for SlopeTree is available at: http://prodata.swmed.edu/download/pub/slopetree_v1/slopetree.tar.gz. Due to their lack of distinct morphological features, bacteria and archaea were extremely difficult to classify until technology was developed to obtain their DNA sequences; these sequences could then be compared to estimate evolutionary relationships. Now, due to technological advances, there is a flood of available sequences from a wide variety of organisms. These advances have spurred the development of algorithms which can estimate evolutionary relationships using whole genomes, in contrast to the more traditional methods which used single genes earlier and now typically use groups of conserved genes. However, there are many challenges when attempting to infer evolutionary relationships, in particular horizontal gene transfer, where DNA is transferred from one organism to another, resulting in an organism’s genome containing DNA that does not reflect its evolution by descent. We developed a new whole-genome method for estimating evolutionary distances which identifies and corrects for horizontal transfer. We found that for SlopeTree and all other whole-genome methods we applied, horizontal transfer causes some evolutionary distances to be grossly underestimated, and that our correction corrects for this.