Effects of nucleotide sequence alignment on phylogeny estimation: A case study of 18S rDNAs of Apicomplexa

Effects of nucleotide sequence alignment on phylogeny estimation: A case study of 18S rDNAs of Apicomplexa
复制标题

DOI:
10.1093/oxfordjournals.molbev.a025779
复制
发表时间:
1997-04-01
影响因子:
10.7
通讯作者:
Ellis, JT
Ellis, JT
中科院分区:
生物学1区
文献类型:
--
作者:
Morrison, DA;Ellis, JT

文献摘要

被引文献

相似文献

系统发育历史的重建是以能够准确地建立性状同源性假设为前提的,这涉及到基于分子序列数据的研究的序列比对。在一项调查核苷酸序列比对的实证研究中,我们基于完整的小亚基 rDNA 序列,使用六种不同的多重比对程序推断了 43 个顶端复合体物种和 3 个恐龙亚目的系统发育树:基于 188 个 rRNA 分子二级结构的手动比对,以及使用 PileUp、ClustalW、TreeAlign、MALIGN 和 SAM 计算机程序的基于相似性的自动比对算法。树是使用邻接法、加权简约法和最大似然法构建的。所有多序列比对程序都产生了相同的基本结构,用于估计类群之间的系统发育关系,这可能代表了潜在的系统发育信号。然而,许多类群的位置对所使用的对齐程序很敏感;不同的排列方式产生的树木平均彼此之间的差异比使用的不同树木构建方法的差异更大。不同程序的多重比对长度差异很大,但比对序列长度并不能很好地预测所得系统发育树的相似性。我们还系统地改变了 ClustalW 程序的空位权重(在序列中插入新空位或扩展已存在空位的相对成本),这产生的比对彼此之间的差异至少与不同比对算法产生的比对相同。此外,尽管许多比对的长度与结构比对的长度相似,但没有间隙权重的组合产生与结构比对相同的树。我们还研究了 rDNA 螺旋和非螺旋区域的系统发育信息内容,并得出结论,螺旋区域信息量最大。因此,我们得出的结论是,许多关于顶复门系统发育的文献分歧可能是基于序列比对策略的差异,而不是数据或树构建方法的差异。
The reconstruction of phylogenetic history is predicated on being able to accurately establish hypotheses of character homology, which involves sequence alignment for studies based on molecular sequence data. In an empirical study investigating nucleotide sequence alignment, we inferred phylogenetic trees for 43 species of the Apicomplexa and 3 of Dinozoa based on complete small-subunit rDNA sequences, using six different multiple-alignment procedures: manual alignment based on the secondary structure of the 188 rRNA molecule, and automated similarity-based alignment algorithms using the PileUp, ClustalW, TreeAlign, MALIGN, and SAM computer programs. Trees were constructed using neighbor-joining, weighted-parsimony, and maximum-likelihood methods. All of the multiple sequence alignment procedures yielded the same basic structure for the estimate of the phylogenetic relationship among the taxa, which presumably represents the underlying phylogenetic signal. However, the placement of many of the taxa was sensitive to the alignment procedure used; and the different alignments produced trees that were on average more dissimilar from each other than did the different tree-building methods used. The multiple alignments from the different procedures varied greatly in length, but aligned sequence length was not a good predictor of the similarity of the resulting phylogenetic trees. We also systematically varied the gap weights (the relative cost of inserting a new gap into a sequence or extending an already-existing gap) for the ClustalW program, and this produced alignments that were at least as different from each other as those produced by the different alignment algorithms. Furthermore, there was no combination of gap weights that produced the same tree as that from the structure alignment, in spite of the fact that many of the alignments were similar in length to the structure alignment. We also investigated the phylogenetic information content of the helical and nonhelical regions of the rDNA, and conclude that the helical regions are the most informative. We therefore conclude that many of the literature disagreements concerning the phylogeny of the Apicomplexa are probably based on differences in sequence alignment strategies rather than differences in data or tree-building methods.