Sequence analysis of the genome of the unicellular cyanobacterium Synechocystis sp. strain PCC6803. II. Sequence determination of the entire genome and assignment of potential protein-coding regions (supplement).

Sequence analysis of the genome of the unicellular cyanobacterium Synechocystis sp. strain PCC6803. II. Sequence determination of the entire genome and assignment of potential protein-coding regions (supplement).
复制标题

DOI:
10.1093/dnares/3.3.185
复制
发表时间:
1996-06-30
期刊:
DNA research : an international journal for rapid publication of reports on genes and genomes
影响因子:
--
通讯作者:
Tabata, S
Tabata, S
中科院分区:
其他
文献类型:
--
作者:
Kaneko, T;Sato, S;Tabata, S

文献摘要

被引文献

相似文献

集胞藻(Synechocystissp.)菌株PCC 6803完成。最终确认的基因组总长度为3,573,470 bp,包括先前报道的1,003,450 bp的序列,从基因组的64%到92%的图谱位置。整个序列由粘粒克隆和λ克隆的基于物理图的重叠群的序列和用于缺口填充的长PCR产物组装而成。通过对整个基因组的两条DNA链进行分析,保证了序列的准确性。通过长PCR产物的限制性分析支持组装序列的真实性,所述长PCR产物使用组装序列数据从基因组DNA直接扩增。为了预测潜在的蛋白质编码区,进行了开放阅读框架(ORF)分析、GeneMark程序分析和数据库相似性搜索。结果表明,在基因组上共分配了3,168个潜在的蛋白质基因,其中145个(4.6%)与报道的基因相同,1,257个(39.6%)和340个(10.8%)分别与报道和假设的基因相似。其余1,426个(45.0%)与数据库中的任何基因都没有明显的相似性。在所分配的潜在蛋白质基因中,有128个与参与光合反应的基因相关。编码潜在蛋白质基因的序列之和占基因组长度的87%。因此,通过添加rRNA和tRNA基因,基因组具有非常紧凑的蛋白质和RNA编码区排列。基因组的一个显著特征是,在整个基因组中发现了99个与转座酶基因相似的开放阅读框,可分为6组,其中至少有26个开放阅读框保持完整。结果表明,该物种在建立期间和建立后经常发生基因组重排。
The sequence determination of the entire genome of theSynechocystissp. strain PCC6803 was completed. The total length of the genome finally confirmed was 3,573,470 bp, including the previously reported sequence of 1,003,450 bp from map position 64% to 92% of the genome. The entire sequence was assembled from the sequences of the physical map-based contigs of cosmid clones and of λ clones and long PCR products which were used for gap-filling. The accuracy of the sequence was guaranteed by analysis of both strands of DNA through the entire genome. The authenticity of the assembled sequence was supported by restriction analysis of long PCR products, which were directly amplified from the genomic DNA using the assembled sequence data. To predict the potential protein-coding regions, analysis of open reading frames (ORFs), analysis by the GeneMark program and similarity search to databases were performed. As a result, a total of 3,168 potential protein genes were assigned on the genome, in which 145 (4.6%) were identical to reported genes and 1,257 (39.6%) and 340 (10.8%) showed similarity to reported and hypothetical genes, respectively. The remaining 1,426 (45.0%) had no apparent similarity to any genes in databases. Among the potential protein genes assigned, 128 were related to the genes participating in photosynthetic reactions. The sum of the sequences coding for potential protein genes occupies 87% of the genome length. By adding rRNA and tRNA genes, therefore, the genome has a very compact arrangement of protein- and RNA-coding regions. A notable feature on the gene organization of the genome was that 99 ORFs, which showed similarity to transposase genes and could be classified into 6 groups, were found spread all over the genome, and at least 26 of them appeared to remain intact. The result implies that rearrangement of the genome occurred frequently during and after establishment of this species.