Protein identification with N and C-terminal sequence tags in proteome projects.

Protein identification with N and C-terminal sequence tags in proteome projects.
复制标题

DOI:
10.1006/jmbi.1998.1726
复制
发表时间:
1998-05
影响因子:
5.6
通讯作者:
M. R. Wilkins;M. R. Wilkins;E. Gasteiger;L. Tonella;Keli Ou;Margaret I. Tyler;Jean‐Charles Sanchez;Andrew A. Gooley;Bradley J. Walsh;A. Bairoch;R. D. Appel;Keith L. Williams;D. F. Hochstrasser;D. F. Hochstrasser
M. R. Wilkins;M. R. Wilkins;E. Gasteiger;L. Tonella;Keli Ou;Margaret I. Tyler;Jean‐Charles Sanchez;Andrew A. Gooley;Bradley J. Walsh;A. Bairoch;R. D. Appel;Keith L. Williams;D. F. Hochstrasser;D. F. Hochstrasser
中科院分区:
生物学2区
文献类型:
--
作者:
M. R. Wilkins;M. R. Wilkins;E. Gasteiger;L. Tonella;Keli Ou;Margaret I. Tyler;Jean‐Charles Sanchez;Andrew A. Gooley;Bradley J. Walsh;A. Bairoch;R. D. Appel;Keith L. Williams;D. F. Hochstrasser;D. F. Hochstrasser

文献摘要

被引文献

相似文献

基因组序列可用于越来越多的生物体。许多这样的生物体的蛋白质组(由基因组表达的蛋白质补体)正在用二维(2D)凝胶电泳进行研究。在这里,我们已经研究了应用短的N-末端和C-末端序列标签的蛋白质分离的2D凝胶上的鉴定。分析了15519个蛋白质的理论N和C末端,这些蛋白质代表了生物体生殖支原体、枯草芽孢杆菌、大肠杆菌、酿酒酵母和人的所有SWISS-PROT条目。发现序列标签具有令人惊讶的特异性,发现43%至83%的蛋白质具有四个氨基酸残基的N末端标签,而74%至97%的蛋白质具有四个氨基酸残基的C末端标签,具体取决于研究的物种。发现五个氨基酸残基的序列标签甚至更特异。为了利用序列标签的这种特异性进行蛋白质鉴定,我们创建了一个全球网络可访问的蛋白质鉴定程序TagIdent(http://www.expasy.ch/www/tools.html),该程序将多达6个氨基酸残基的序列标签以及估计的蛋白质pI和质量与SWISS-PROT数据库中的蛋白质进行匹配。我们证明了这种识别方法的实用性与91个不同的大肠杆菌产生的序列标签。大肠杆菌蛋白质通过二维凝胶电泳纯化。51个蛋白质被明确地确定凭借其序列标签和估计的pI和质量,并进一步确定11个蛋白质时,序列标签与蛋白质氨基酸组成数据相结合。我们的结论是,TagIdent鉴定方法是最适合于从原核生物的完整基因组序列的蛋白质的鉴定。该方法不太适合来自真核生物的蛋白质,因为许多真核生物蛋白质不适合通过Edman降解进行测序,并且除非生物体的完整序列可用,否则标签蛋白质鉴定不能明确。
Genome sequences are available for increasing numbers of organisms. The proteomes (protein complement expressed by the genome) of many such organisms are being studied with two-dimensional (2D) gel electrophoresis. Here we have investigated the application of short N-terminal and C-terminal sequence tags to the identification of proteins separated on 2D gels. The theoretical N and C termini of 15, 519 proteins, representing all SWISS-PROT entries for the organisms Mycoplasma genitalium, Bacillus subtilis, Escherichia coli, Saccharomyces cerevisiae and human, were analysed. Sequence tags were found to be surprisingly specific, with N-terminal tags of four amino acid residues found to be unique for between 43% and 83% of proteins, and C-terminal tags of four amino acid residues unique for between 74% and 97% of proteins, depending on the species studied. Sequence tags of five amino acid residues were found to be even more specific. To utilise this specificity of sequence tags for protein identification, we created a world-wide web-accessible protein identification program, TagIdent (http://www.expasy.ch/www/tools.html), which matches sequence tags of up to six amino acid residues as well as estimated protein pI and mass against proteins in the SWISS-PROT database. We demonstrate the utility of this identification approach with sequence tags generated from 91 different E. coli proteins purified by 2D gel electrophoresis. Fifty-one proteins were unambiguously identified by virtue of their sequence tags and estimated pI and mass, and a further 11 proteins identified when sequence tags were combined with protein amino acid composition data. We conlcude that the TagIdent identification approach is best suited to the identification of proteins from prokaryotes whose complete genome sequences are available. The approach is less well suited to proteins from eukaryotes, as many eukaryotic proteins are not amenable to sequencing via Edman degradation, and tag protein identification cannot be unambiguous unless an organism's complete sequence is available.