A lineage-specific gene encoding a major matrix protein of the sea urchin embryo spicule. II. Structure of the gene and derived sequence of the protein.

A lineage-specific gene encoding a major matrix protein of the sea urchin embryo spicule. II. Structure of the gene and derived sequence of the protein.
复制标题

编码海胆胚胎骨刺主要基质蛋白的谱系特异性基因。

DOI:
10.1016/0012-1606(87)90254-5
复制
发表时间:
1987
影响因子:
2.7
通讯作者:
Davidson,EH
Davidson,EH
中科院分区:
生物学3区
文献类型:
--
作者:
Sucov,HM;Benson,S;Robinson,JJ;Britten,RJ;Wilt,F;Davidson,EH

文献摘要

被引文献

相似文献

用多克隆抗骨针基质蛋白抗血清分离的λ gt 11 cDNA克隆见所附论文[S. C. Benson,H. M.苏科夫湖Stephens,E. H. Davidson和F. 03 The Dog(1987)120,499 -506]以编码突出的50-kDa骨针基质蛋白(SM 50)。用该克隆筛选同源重组子,并测定基因结构。SM 50基因在每个单倍体基因组中出现一次。它含有一个位于第35位密码子内的内含子。通过引物延伸定位了翻译起始信号前110个核苷酸对的独特转录起始位点。该基因全长1895个核苷酸,不包括3′端poly(A)序列,含有一个450个密码子的开放阅读框架。尽管SM 50 mRNA在整个胚胎RNA中很少见,但据计算,SM 50 mRNA的普遍性约占成骨间充质细胞中总mRNA的1%。推导的肽序列表明其N端为典型的信号肽,C端附近有一个N-连接的糖基化位点。该蛋白约45%的长度包含在一个由13个氨基酸的连续近似重复组成的结构域中,其共有序列是。该蛋白还包含一个异常富含脯氨酸残基的内部结构域和一个非常碱性的C-末端区域。
A λgt11 cDNA clone isolated by use of a polyclonal antispicule matrix protein antiserum is shown in the accompanying paper [S. C. Benson, H. M. Sucov, L. Stephens, E. H. Davidson, and F. Wilt (1987)Dev. Biol.120,499–506] to encode a prominent 50-kDa spicule matrix protein (SM50). This clone was used to select homologous genomic recombinants, and the structure of the gene was determined. The SM50 gene occurs once per haploid genome. It contains a single intron located within the 35th codon. A unique transcription initiation site 110 nucleotide pairs prior to the translation start signal was mapped by primer extension. The mRNA is 1895 nucleotides in length, excluding the 3′ poly(A) sequence, and contains a single open reading frame 450 codons in length. Though rare in whole embryo RNA the prevalence of the SM50 mRNA is calculated to be about 1% of the total mRNA in skeletogenic mesenchyme cells. The derived peptide sequence indicates a typicalN-terminal signal peptide, and anN-linked glycosylation site near theCterminus. About 45% of the length of the protein is included in a domain composed of consecutive approximate repetitions of a 13-aminoacid element, the consensus sequence of which isThe protein also contains an internal domain unusually rich in proline residues and a very basicC-terminal region.