Searching for RNA genes using base-composition statistics

Searching for RNA genes using base-composition statistics
复制标题

DOI:
10.1093/nar/30.9.2076
复制
发表时间:
2002-05-01
影响因子:
14.9
通讯作者:
Schattner, P
Schattner, P
中科院分区:
生物学2区
文献类型:
--
作者:
Schattner, P

文献摘要

被引文献

相似文献

已经研究了这样的假设,即富含非蛋白质编码RNA(NcRNAs)的基因组区域可以利用单碱基和二核苷酸统计中的局部变异来识别。比较了7类ncRNA和3个基因组的(G+C)%、(G-C)%差、(A-T)%差和二核苷酸频率统计。在(G+C)%和在简氏甲烷球菌中,二核苷酸‘CG’的频率有显著差异。开发了基于这两个碱基组成统计的筛选程序。仅用(G+C)%筛选,就可以鉴定出1%的jannaschii基因组包含所有44个已知的转移RNA、核糖体RNA和信号识别颗粒RNA。当使用(G+C)%结合CG二核苷酸频率筛选时,44个已知的jannaschii结构ncRNA中的43个被再次识别,而与已知或推测的蛋白质编码基因重叠的可能错误命中的数量从15个减少到6个。此外,还鉴定了19个候选ncRNA,其中一个与几个已知的古生代RNaseP RNA有显著同源性。
The hypothesis that genomic regions rich in non-protein-coding RNAs (ncRNAs) can be identified using local variations in single-base and dinucleotide statistics has been investigated. (G+C)%, (G-C)% difference, (A-T)% difference and dinucleotide-frequency statistics were compared among seven classes of ncRNAs and three genomes. Significant variations were observed in (G+C)% and, in Methanococcus jannaschii, in the frequency of the dinucleotide 'CG'. Screening programs based on these two base-composition statistics were developed. With (G+C)% screening alone, a 1% fraction of the M.jannaschii genome containing all 44 known transfer RNAs, ribosomal RNAs and signal recognition particle RNAs could be identified. When (G+C)% combined with CG dinucleotide-frequency screening was used, 43 of the 44 known M.jannaschii structural ncRNAs were again identified, while the number of presumably false hits overlapping a known or putative protein-coding gene was reduced from 15 to 6. In addition, 19 candidate ncRNAs were identified including one with significant homology to several known archaeal RNaseP RNAs.