Identification of shared single copy nuclear genes in Arabidopsis, Populus, Vitis and Oryza and their phylogenetic utility across various taxonomic levels

Identification of shared single copy nuclear genes in Arabidopsis, Populus, Vitis and Oryza and their phylogenetic utility across various taxonomic levels
复制标题

DOI:
10.1186/1471-2148-10-61
复制
发表时间:
2010-02-24
影响因子:
3.4
通讯作者:
dePamphilis, Claude W.
dePamphilis, Claude W.
中科院分区:
生物学2区
文献类型:
--
作者:
Duarte, Jill M.;Wall, P. Kerr;dePamphilis, Claude W.

文献摘要

被引文献

相似文献

背景:尽管在被子植物中发现的绝大多数基因是基因家族的成员,且基因和基因组重复在植物基因组中是普遍存在的力量,但一些基因与基因组中的所有其他基因有足够大的差异,以至于它们在操作上可被定义为“单拷贝”。利用基因聚类算法MCL - tribe,我们已经确定了一组959个单拷贝基因,它们是拟南芥(Arabidopsis thaliana)、毛果杨(Populus trichocarpa)、葡萄(Vitis vinifera)和水稻(Oryza sativa)基因组中的共享单拷贝基因。为了对这些基因进行表征,我们进行了多项分析,检查基因本体论(GO)注释、编码序列长度、外显子数量、结构域数量、在远缘谱系(如卷柏属(Selaginella)和小立碗藓属(Physcomitrella))中的存在情况,以及系统发育分析以估计在其他种子植物中的拷贝数并证明它们在系统发育方面的用途。然后我们提供了一些例子,说明如何在系统发育分析中利用这些基因来重建生物历史,既可以利用种子植物的EST数据库中的现存覆盖范围,也可以通过在十字花科(Brassicaceae)中进行逆转录聚合酶链反应(RT - PCR)从头扩增。 结果:在拟南芥、毛果杨、葡萄和水稻中共有959个单拷贝核基因[“APVO单拷贝基因”]。这些基因中的大多数也存在于卷柏属和小立碗藓属的基因组中。197个物种的公共EST数据集表明,这些基因中的大多数存在于各种各样的种子植物集合中,并且似乎以单拷贝或极低拷贝基因的形式存在,尽管在近期多倍体类群以及有明显证据表明存在共享大规模重复事件的谱系中存在例外情况。编码定位于细胞器的蛋白质的基因比随机预期更常为单拷贝,但导致这种偏向的进化力量尚不清楚。无论在不同的开花植物谱系中导致大量共享单拷贝基因的进化机制如何,这些基因对于系统发育和比较分析都是有价值的。使用RT - PCR在十字花科中扩增了18个APVO单拷贝基因并直接测序。与最近使用质体和ITS序列的研究相比,这些序列的比对为十字花科系统发育提供了更高的分辨率。对主要来自公共EST数据库的69种种子植物的13个APVO单拷贝基因的序列分析产生了一个系统发育树,该树在很大程度上与基于多个质体序列的先前假设一致。由于序列信息有限,依赖EST序列的单基因系统发育树的自举支持有限,但串联比对会产生对已确定关系具有强自举支持的系统发育树。总体而言,这些单拷贝核基因是系统发育研究中有希望的标记,并且比来自质体或线粒体基因组的常用蛋白质编码序列包含更高比例的具有系统发育信息的位点。 结论:假定直系同源的共享单拷贝核基因为植物系统发育、基因组图谱绘制和其他应用提供了大量新的证据来源,也是一类需要进行功能表征的大量基因。初步证据表明,本研究中确定的许多共享单拷贝核基因可能非常适合作为解决各种分类水平上的系统发育假设的标记。
Background: Although the overwhelming majority of genes found in angiosperms are members of gene families, and both gene- and genome-duplication are pervasive forces in plant genomes, some genes are sufficiently distinct from all other genes in a genome that they can be operationally defined as 'single copy'. Using the gene clustering algorithm MCL-tribe, we have identified a set of 959 single copy genes that are shared single copy genes in the genomes of Arabidopsis thaliana, Populus trichocarpa, Vitis vinifera and Oryza sativa. To characterize these genes, we have performed a number of analyses examining GO annotations, coding sequence length, number of exons, number of domains, presence in distant lineages, such as Selaginella and Physcomitrella, and phylogenetic analysis to estimate copy number in other seed plants and to demonstrate their phylogenetic utility. We then provide examples of how these genes may be used in phylogenetic analyses to reconstruct organismal history, both by using extant coverage in EST databases for seed plants and de novo amplification via RT-PCR in the family Brassicaceae.Results: There are 959 single copy nuclear genes shared in Arabidopsis, Populus, Vitis and Oryza ["APVO SSC genes"]. The majority of these genes are also present in the Selaginella and Physcomitrella genomes. Public EST sets for 197 species suggest that most of these genes are present across a diverse collection of seed plants, and appear to exist as single or very low copy genes, though exceptions are seen in recently polyploid taxa and in lineages where there is significant evidence for a shared large-scale duplication event. Genes encoding proteins localized in organelles are more commonly single copy than expected by chance, but the evolutionary forces responsible for this bias are unknown.Regardless of the evolutionary mechanisms responsible for the large number of shared single copy genes in diverse flowering plant lineages, these genes are valuable for phylogenetic and comparative analyses. Eighteen of the APVO SSC single copy genes were amplified in the Brassicaceae using RT-PCR and directly sequenced. Alignments of these sequences provide improved resolution of Brassicaceae phylogeny compared to recent studies using plastid and ITS sequences. An analysis of sequences from 13 APVO SSC genes from 69 species of seed plants, derived mainly from public EST databases, yielded a phylogeny that was largely congruent with prior hypotheses based on multiple plastid sequences. Whereas single gene phylogenies that rely on EST sequences have limited bootstrap support as the result of limited sequence information, concatenated alignments result in phylogenetic trees with strong bootstrap support for already established relationships. Overall, these single copy nuclear genes are promising markers for phylogenetics, and contain a greater proportion of phylogenetically-informative sites than commonly used protein-coding sequences from the plastid or mitochondrial genomes.Conclusions: Putatively orthologous, shared single copy nuclear genes provide a vast source of new evidence for plant phylogenetics, genome mapping, and other applications, as well as a substantial class of genes for which functional characterization is needed. Preliminary evidence indicates that many of the shared single copy nuclear genes identified in this study may be well suited as markers for addressing phylogenetic hypotheses at a variety of taxonomic levels.