Frequent oligonucleotides and peptides of the Haemophilus influenzae genome

Frequent oligonucleotides and peptides of the Haemophilus influenzae genome
复制标题

DOI:
10.1093/nar/24.21.4263
复制
发表时间:
1996-11-01
影响因子:
14.9
通讯作者:
Campbell, AM
Campbell, AM
中科院分区:
生物学2区
文献类型:
--
作者:
Karlin, S;Mrazek, J;Campbell, AM

文献摘要

被引文献

相似文献

完整的流感嗜血杆菌基因组(1.83 Mb,Rd菌株)提供了表征全球基因组不均一性和检测重要序列信号的机会。沿着这些路线,用于识别频繁词(寡核苷酸和/或肽)及其分布的新方法被应用于流感嗜血杆菌基因组,并与其他细菌基因组的频繁词进行一些比较和对比。常见的寡核苷酸有三大类:(i)与常见的摄取信号序列(USS)、AAGTGCGGT(USS+)及其反向互补序列(USS-)相关的寡核苷酸,(ii)多个四核苷酸重复和(iii)基因间二联体序列(ISD),发现为AAGCCCACCCTAC及其二联体形式。USS+和USS-以几乎相等的计数出现,在基因组中分布非常均匀,并且主要出现在蛋白质编码结构域的相同阅读框中(USS+翻译为Ser-Ala-瓦尔,USS-翻译为Thr-Ala-Leu)。这些观察结果表明,USS有助于全球基因组功能,例如,在复制和/或修复过程中,或作为膜附着位点,或作为帮助包装DNA的序列。长的四核苷酸迭代,实际上是流感嗜血杆菌独有的(即,在其他原核生物中未知),通过复制和/或同源重组期间的聚合酶滑动,可以产生表达替代蛋白的亚群。13 bp的频繁IDS的话,总是intergenic的,主要发生在集群,并提供了潜在的复杂的二级结构,表明这些序列可能是重要的信号调节其侧翼基因的活性。流感嗜血杆菌的常见寡肽主要有两类--由寡核苷酸频繁词(USS,四核苷酸重复)诱导的寡肽,以及与ATP或GTP结合位点相关的寡肽,这些结合位点通常由三个基序组成:有助于描绘结合口袋的A盒;在水解中起作用的B盒;以及功能未知的C盒。A盒在原核生物和真核生物中相当普遍。B基序和C基序似乎专门用于各种官能团(例如,转运、重组、伴侣活性)。其他推定的基序对应于大肠杆菌基序的同源物,例如,与转录加工蛋白质、氨酰-tRNA合成酶和在电子传递中起作用的蛋白质相关。
The complete Haemophilus influenzae genome (1.83 Mb, Rd strain) provides opportunities for characterizing global genomic inhomogeneities and for detecting important sequence signals. Along these lines, new methods for identifying frequent words (oligonucleotides and/or peptides) and their distributions are applied to the H.influenzae genome with some comparisons and contrasts made with frequent words of other bacterial genomes. Three major classes of frequent oligonucleotides stand out: (i) oligos related to the familiar uptake signal sequences (USSs), AAGTGCGGT (USS+) and its inverted complement (USS-), (ii) multiple tetranucleotide iterations and (iii) intergenic dyad sequences (ISDs) found as AAGCCCACCCTAC and its dyad form. The USS+ and USS- occur in almost equal counts, are remarkably evenly spaced around the genome, and appear predominantly in the same reading frame of protein coding domains (USS+ translated to Ser-Ala-Val, USS- translated to Thr-Ala-Leu). These observations suggest that USSs contribute to global genomic functions, for example, in replication and/or repair processes, or as membrane attachment sites, or as sequences helping to pack DNA. The long tetranucleotide iterations, virtually unique to H.influenzae (i.e., unknown in other prokaryotes), through polymerase slippage during replication and/or homologous recombination may produce subpopulations expressing alternative proteins. The 13 bp frequent IDS words, invariably intergenic, occur mostly in clusters and provide potential for complex secondary structures suggesting that these sequences may be important signals for regulating the activity of their flanking genes. The frequent oligopeptides of H.influenzae are principally of two kinds-those induced by oligonucleotide frequent words (USSs, tetranucleotide iterations), and those associated with ATP or GTP binding sites that are generally composed of three motifs: the A-box which contributes to delineating the binding pocket; the B-box which functions in hydrolysis; and the C-box whose function is unknown. The A-box occurs fairly universally in prokaryotes and eukaryotes. The B- and C-motifs appear to be specialized to various functional groups (e.g., transport, recombination, chaperone activity). Other putative motifs correspond to homologs of Escherichia coli motifs, for example, are associated with proteins of transcriptional processing, aminoacyl-tRNA synthetases and proteins functioning in electron transfer.