On conditioned reconstruction, gene content data, and the recovery of fusion genomes

On conditioned reconstruction, gene content data, and the recovery of fusion genomes
复制标题

DOI:
10.1016/j.ympev.2005.11.020
复制
发表时间:
2006-04-01
影响因子:
4.1
通讯作者:
Houde, P
Houde, P
中科院分区:
生物学1区
文献类型:
--
作者:
Bailey, CD;Fain, MG;Houde, P

文献摘要

被引文献

相似文献

条件重建(CR)1代表了一种新的系统发育方法,其已被提出作为利用大量基因缺失/存在数据来重建系统发育关系并直接研究基因组融合对进化的影响的手段(Lake和里维拉,2004;里维拉和Lake,2004; Simonson等人,2005年)。在CR的第一次直接应用中,结果明确支持古细菌和真细菌基因组融合事件在真核基因组祖先中的作用(McInerney和威尔金森,2005;里维拉和莱克,2004;西蒙森等人,2005年)。在本手稿中,我们概述了CR的基本组成部分,讨论了其应用背后的逻辑,并随后确定了数据解释方面的问题,以及目前在基因组融合研究中使用条件基因组和对齐网络的问题。条件重建开始于选择条件基因组(CG),其将代表在基质发育期间编码的全套直链淀粉基因(Lake和里维拉,2004)。一直强调严格使用相对较小的基因组作为条件基因组(Lake和里维拉,2004)。这一建议是基于这样的观察:使用相对较大的条件基因组可以诱导一种称为“大基因组吸引力”的假象,这可能会误导系统发育推断。将用于分析的所有其他基因组针对CG中每个基因的推定直系同源物的存在或不存在进行评分。不考虑分析中包括的其他基因组中存在但不存在于CG中的直系同源物,并且将CG从系统发育分析中排除。实施这种方法是为了简化基因选择过程,因为在分析中纳入CG将不允许缺席!缺席(A! A)待定义的变换(Lake and里维拉,2004),这代表了Lake and里维拉(2004)所青睐的新条件成对距离的估计问题。因此,有人认为,使用条件基因组允许唯一定义的概率的所有字符状态转换(例如,两个分类群A! A,A! P,P! A,P!通过该过程开发的矩阵(在此称为"条件矩阵")经受多次自举重采样和系统发育分析,以识别由数据支持的最优和次优拓扑。虽然各种方法都适用于条件矩阵的分析(例如,简约性和简单距离),但Lake和里维拉(2004)得出结论,一种用于估计基因插入/缺失的条件概率的基于马尔可夫的新方法,需要开发条件矩阵,产生的结果最不容易受到"大基因组吸引"的影响(下文进一步讨论)。在系统发育分析之后,尝试将从分析中恢复的所有网络进行比对(参见Lake和里维拉,2004中的图3)。网络对齐是网络围绕节点或反转网络的成功旋转,使得来自两个或更多个不同网络的操作终端可以以重复的线性模式部分或完全重叠。对齐的网络被认为是网状进化的证据,而不是严格的分歧。使用替代单个CG重复该过程以及持续恢复与相同融合假设一致的结果被接受为融合的强有力证据(里维拉和莱克,2004年-补充材料)。
Conditioned reconstruction (CR) 1 represents a new phylogenetic method that has been presented as a means of utilizing vast amounts of gene absence/presence data to reconstruct phylogenetic relationships and to directly study the influence of genome fusion on evolution (Lake and Rivera, 2004; Rivera and Lake, 2004; Simonson etal., 2005). In the first direct application of CR, the results were stated to unambiguously support the role of an archaeal and eubacterial genome fusion event in the ancestry of the eukaryotic genome (McInerney and Wilkinson, 2005; Rivera and Lake, 2004; Simonson et al., 2005). In the present manuscript, we outline the basic components of CR, discuss the logic behind their application, and subsequently identify concerns with aspects of data interpretation and the current use of conditioning genomes and aligned networks in the study of genome fusion. Conditioned reconstruction begins with the selection of a conditioning genome (CG) that will represent the full set of orthologous genes coded during matrix development (Lake and Rivera, 2004). Strong emphasis has been placed on the strict use of relatively small genomes as conditioning genomes (Lake and Rivera, 2004). This recommendation is based on the observation that the use of relatively large conditioning genomes can induce an artifact referred to as “big genome attraction,” which can mislead phylogenetic inference. All other genomes that will be used in the analysis are scored for the absence or presence of a putative ortholog to each gene in the CG. Orthologs present in other genomes included in the analysis, but absent from the CG, are not considered and the CG is excluded from phylogenetic analysis. This approach was implemented to simplify the gene selection process and because the inclusion of the CG in analyses would not allow Absence! Absence (A! A) transformations to be defined (Lake and Rivera, 2004), which represents a problem for the estimation of novel conditioned pairwise distances favored by Lake and Rivera (2004). Thus, it was argued that the use of a conditioning genome allows for uniquely defined probabilities of all character state transformations (eg, for two taxa A! A, A! P, P! A, P! P).The matrix developed through this procedure (referred to here as the “conditioned matrix”) is subject to multiple bootstrap resamplings and phylogenetic analysis to identify both optimal and suboptimal topologies supported by the data. While various approaches are applicable in the analysis of the conditioned matrix (eg, parsimony and simple distance), Lake and Rivera (2004) concluded that a novel Markov-based approach used to estimate conditional probabilities of gene insertion/deletion, and that required the development of the conditioned matrix, produced results that were least susceptible to “big genome attraction”(additional discussion below). Subsequent to phylogenetic analysis an attempt is made to align all networks recovered from the analysis (see Fig. 3 in Lake and Rivera, 2004). Network alignment is the successful rotating of networks around nodes or inverting networks such that the operational terminals from two or more different networks may be partially or fully overlapped in a repeat linear pattern. Aligned networks are accepted as evidence of reticulate evolution rather than strict divergence. Repetition of this process using alternative single CGs and the continued recovery of results that are consistent with the same fusion hypothesis are accepted as robust evidence of fusion (Rivera and Lake, 2004—supplemental material).