COMPLETE NUCLEOTIDE-SEQUENCE OF BACTERIOPHAGE-T7 DNA AND THE LOCATIONS OF T7 GENETIC ELEMENTS

COMPLETE NUCLEOTIDE-SEQUENCE OF BACTERIOPHAGE-T7 DNA AND THE LOCATIONS OF T7 GENETIC ELEMENTS
复制标题

DOI:
10.1016/s0022-2836(83)80282-4
复制
发表时间:
1983-01-01
影响因子:
5.6
通讯作者:
STUDIER, FW
STUDIER, FW
中科院分区:
生物学2区
文献类型:
--
作者:
DUNN, JJ;STUDIER, FW

文献摘要

被引文献

相似文献

用Maxam&Gilbert技术测定了T7噬菌体DNA的39,936个碱基对的全序列。所有先前已知的T7基因和几个未被怀疑的基因都已在该序列中被识别出来。T7DNA携带遗传信息的效率很高:50个基因的编码序列紧密排列,但基本上不重叠,几乎占据了核苷酸序列的92%。这种排列强烈地表明,所有50个紧密结合的基因都得到了表达,尽管到目前为止只有38个基因表达的证据。此外,已经确定了五个潜在的重叠基因,并有初步证据表明其中一个基因表达了。在发现编码序列之间的间隙时,它们通常不到100个碱基对,并且通常包含一个或多个转录信号、RNAase III裂解位点或复制起点。T7DNA中的转录信号包括3个强烈的早期启动子和大肠杆菌RNA聚合酶的早期终止位点,以及17个启动子和1个终止位点。已经定位了10个RNAase III裂解位点,其中5个在早期区域,5个在晚期区域。初级转录本在这些位点被处理,以提供在体内观察到的信使RNA。几乎所有的T7信使RNA都是多顺反子的,但在转录或翻译水平上几乎没有极性效应,而且大多数T7蛋白似乎是独立启动的,每个蛋白都有自己的核糖体结合和起始点。大多数T7蛋白的起始密码子都是AUG,但也有少数蛋白质的起始密码子是GAG。某些T7基因规定了一对重叠的蛋白质。基因4指定的两种蛋白质的组成大致相同,起始于同一阅读框中两个不同的核糖体结合和起始位置,结束于共同的终止密码子。基因10所指定的两种蛋白质的含量非常不同。它们在相同的起始点开始,但次要基因10蛋白似乎是通过翻译阅读框在正常终止密码子之前的移位而产生的,从而在主要蛋白的COOH末端增加了53个氨基酸。基因10指定了噬菌体颗粒的主要衣壳蛋白,主要和次要基因10蛋白都被整合到噬菌体颗粒中。另外一个或两个T7基因似乎利用翻译移码来产生不同数量的蛋白质,这些蛋白质在其COOH末端不同。给出了预测的所有T7蛋白的氨基酸序列和组成(移码产生的蛋白除外)。T7DNA以160个碱基对的完美直接重复开始和结束。紧邻这个末端重复,在成熟DNA的两端,有非常相似的,规则的排列,由一个七个碱基序列的12个不完整的拷贝组成。这些阵列占据了大约160个碱基对,从末端重复开始大约15个碱基对。在T7DNA的串联形式中,末端重复的单个副本被这两个重复序列阵列所包围,这种排列似乎以某种方式参与了成熟T7DNA末端的形成。
The complete nucleotide sequence of bacteriophage T7 DNA, 39,936 base-pairs, has been determined by the techniques of Maxam & Gilbert. All previously known T7 genes and several unsuspected genes have been identified in the sequence. T7 DNA carries genetic information very efficiently: the coding sequences of 50 genes are close-packed but essentially not overlapping, and occupy almost 92% of the nucleotide sequence. This arrangement strongly suggests that all 50 of these closepacked genes are expressed, although there is as yet evidence for expression of only 38 of them. In addition, five potential overlapping genes have been identified, and there is preliminary evidence that one of them is expressed. Where gaps between coding sequences are found, they usually are less than 100 basepairs long, and usually contain one or more transcription signals, RNAase III cleavage sites, or origins of replication. Transcription signals in the T7 DNA include the three strong early promoters and the early termination site forEscherichia coliRNA polymerase, and 17 promoters and one termination site for T7 RNA polymerase. Ten RNAase III cleavage sites have been located, five in the early region and five in the late region. The primary transcripts are processed at these sites to provide the messenger RNAs observedin vivo. Almost all of the T7 messenger RNAs are polycistronic, but there are few polar effects at the level of transcription or translation, and most T7 proteins seem to be initiated independently, each from its own ribosome-binding and initiation site. The initiation codon for most T7 proteins is AUG, but a few proteins are predicted to begin at GUG. Certain T7 genes specify pairs of overlapping proteins. The two proteins specified by gene 4 are made in about equal amounts, beginning at two different ribosome-binding and initiation sites in the same reading frame and ending at a common termination codon. The two proteins specified by gene 10 are made in very different amounts. They begin at the same initiation site, but the minor gene 10 protein appears to be produced by a shift in translational reading frame just ahead of the normal termination codon, thereby adding 53 amino acids to the COOH-terminal end of the major protein. Gene 10 specifies the major capsid protein of the phage particle, and both the major and minor gene 10 proteins are incorporated into the phage particle. One or two other T7 genes appear to utilize translational frameshifting to produce unequal amounts of proteins that differ at their COOH-terminal ends. The amino acid sequences and compositions predicted for all of the T7 proteins (except the proteins produced by frameshifting) are given. T7 DNA begins and ends with a perfect direct repeat of 160 base-pairs. Immediately adjacent to this terminal repetition, at both ends of the mature DNA, lie very similar, regular arrays of 12 imperfect copies of a seven-base sequence. These arrays occupy about 160 base-pairs, starting about 15 basepairs from the terminal repetition. In the concatemeric form of T7 DNA, a single copy of the terminal repetition is flanked by these two arrays of repeated sequences, and it seems likely that this arrangement is involved somehow in formation of the ends of mature T7 DNA.