Genome-wide evolutionary analysis of the noncoding RNA genes and noncoding DNA of Paramecium tetraurelia

Genome-wide evolutionary analysis of the noncoding RNA genes and noncoding DNA of Paramecium tetraurelia
复制标题

四尿草履虫非编码 RNA 基因和非编码 DNA 的全基因组进化分析。

DOI:
10.1261/rna.1306009
复制
发表时间:
2009-04-01
期刊:
RNA
影响因子:
4.5
通讯作者:
Amar, Laurence
Amar, Laurence
中科院分区:
生物学3区
文献类型:
--
作者:
Chen, Chun-Long;Zhou, Hui;Amar, Laurence

文献摘要

被引文献

相似文献

单细胞真核生物四脲草履虫(Parameciumtetraurelia)的基因组包含非编码DNA(ncDNA),平均基因间序列数超过39,000,内含子数超过90,000,平均长度分别为390 bp和25 bp。在这里,我们分析了该基因组的ncRNA基因、内含子和基因间序列的分子特征。我们主要使用计算程序和比较基因组学,因为P. tetraurelia基因组已经在整个全基因组复制(WGD)中形成。我们鉴定了417个5S rRNA、snRNA、snoRNA、SRP RNA和tRNA推定基因,其中415个在基因间序列内,2个在内含子内。这些ncRNA基因的进化似乎主要涉及纯化选择和基因缺失。然后,我们比较了中断最近WGD中出现的蛋白质编码基因重复的内含子,并鉴定了在最严格的限制下进化的数千个内含子的群体(>95%的同一性)。我们还表明,低核苷酸取代水平的特点,分别为50和80-115碱基对侧翼,终止和起始密码子的蛋白质编码基因。较低的取代水平标记高度转录基因侧翼的碱基对,或具有大量WGD相关序列的基因组的起始密码子。最后,毗邻蛋白质编码基因,我们的特点是32个DNA基序能够编码稳定和进化保守的RNA二级结构,并定义推定的表达控制元件。14个具有相似性质的DNA基序远离蛋白质编码基因,并可能编码调控ncRNA。
The compact genome of the unicellular eukaryote Paramecium tetraurelia contains noncoding DNA (ncDNA) distributed into >39,000 intergenic sequences and >90,000 introns of 390 base pairs (bp) and 25 bp on average, respectively. Here we analyzed the molecular features of the ncRNA genes, introns, and intergenic sequences of this genome. We mainly used computational programs and comparative genomics possible because the P. tetraurelia genome had formed throughout whole-genome duplications (WGDs). We characterized 417 5S rRNA, snRNA, snoRNA, SRP RNA, and tRNA putative genes, 415 of which map within intergenic sequences, and two, within introns. The evolution of these ncRNA genes appears to have mainly involved purifying selection and gene deletion. We then compared the introns that interrupt the protein-coding gene duplicates arisen from the recent WGD and identified a population of a few thousands of introns having evolved under most stringent constraints (>95% of identity). We also showed that low nucleotide substitution levels characterize the 50 and 80-115 base pairs flanking, respectively, the stop and start codons of the protein-coding genes. Lower substitution levels mark the base pairs flanking the highly transcribed genes, or the start codons of the genes of the sets with a high number of WGD-related sequences. Finally, adjacent to protein-coding genes, we characterized 32 DNA motifs able to encode stable and evolutionary conserved RNA secondary structures and defining putative expression controlling elements. Fourteen DNA motifs with similar properties map distant from protein-coding genes and may encode regulatory ncRNAs.