On conditioned reconstruction, gene content data, and the recovery of fusion genomes
On conditioned reconstruction, gene content data, and the recovery of fusion genomes
复制标题
DOI:
10.1016/j.ympev.2005.11.020
复制
发表时间:
2006-04-01
影响因子:
4.1
通讯作者:
Houde, P
中科院分区:
文献类型:
--
作者:
Bailey, CD;Fain, MG;Houde, P
Conditioned reconstruction (CR) 1 represents a new phylogenetic method that has been presented as a means of utilizing vast amounts of gene absence/presence data to reconstruct phylogenetic relationships and to directly study the influence of genome fusion on evolution (Lake and Rivera, 2004; Rivera and Lake, 2004; Simonson etal., 2005). In the first direct application of CR, the results were stated to unambiguously support the role of an archaeal and eubacterial genome fusion event in the ancestry of the eukaryotic genome (McInerney and Wilkinson, 2005; Rivera and Lake, 2004; Simonson et al., 2005). In the present manuscript, we outline the basic components of CR, discuss the logic behind their application, and subsequently identify concerns with aspects of data interpretation and the current use of conditioning genomes and aligned networks in the study of genome fusion. Conditioned reconstruction begins with the selection of a conditioning genome (CG) that will represent the full set of orthologous genes coded during matrix development (Lake and Rivera, 2004). Strong emphasis has been placed on the strict use of relatively small genomes as conditioning genomes (Lake and Rivera, 2004). This recommendation is based on the observation that the use of relatively large conditioning genomes can induce an artifact referred to as “big genome attraction,” which can mislead phylogenetic inference. All other genomes that will be used in the analysis are scored for the absence or presence of a putative ortholog to each gene in the CG. Orthologs present in other genomes included in the analysis, but absent from the CG, are not considered and the CG is excluded from phylogenetic analysis. This approach was implemented to simplify the gene selection process and because the inclusion of the CG in analyses would not allow Absence! Absence (A! A) transformations to be defined (Lake and Rivera, 2004), which represents a problem for the estimation of novel conditioned pairwise distances favored by Lake and Rivera (2004). Thus, it was argued that the use of a conditioning genome allows for uniquely defined probabilities of all character state transformations (eg, for two taxa A! A, A! P, P! A, P! P).The matrix developed through this procedure (referred to here as the “conditioned matrix”) is subject to multiple bootstrap resamplings and phylogenetic analysis to identify both optimal and suboptimal topologies supported by the data. While various approaches are applicable in the analysis of the conditioned matrix (eg, parsimony and simple distance), Lake and Rivera (2004) concluded that a novel Markov-based approach used to estimate conditional probabilities of gene insertion/deletion, and that required the development of the conditioned matrix, produced results that were least susceptible to “big genome attraction”(additional discussion below). Subsequent to phylogenetic analysis an attempt is made to align all networks recovered from the analysis (see Fig. 3 in Lake and Rivera, 2004). Network alignment is the successful rotating of networks around nodes or inverting networks such that the operational terminals from two or more different networks may be partially or fully overlapped in a repeat linear pattern. Aligned networks are accepted as evidence of reticulate evolution rather than strict divergence. Repetition of this process using alternative single CGs and the continued recovery of results that are consistent with the same fusion hypothesis are accepted as robust evidence of fusion (Rivera and Lake, 2004—supplemental material).