THE PROSITE DICTIONARY OF SITES AND PATTERNS IN PROTEINS, ITS CURRENT STATUS

THE PROSITE DICTIONARY OF SITES AND PATTERNS IN PROTEINS, ITS CURRENT STATUS
复制标题

DOI:
10.1093/nar/21.13.3097
复制
发表时间:
1993-07-01
影响因子:
14.9
通讯作者:
BAIROCH, A
BAIROCH, A
中科院分区:
生物学2区
文献类型:
--
作者:
BAIROCH, A

文献摘要

被引文献

相似文献

背景ProSite是蛋白质序列中发现的位点和模式的汇编;它可以作为一种方法来确定从基因组或cDNA序列翻译而来的未鉴定蛋白质的功能。在一些情况下,未知蛋白质的序列与已知结构的任何蛋白质的关系太远,不能通过整体序列比对来检测其相似性,但可以通过在其序列中出现特定的残基类型簇来揭示关系,该特定残基类型簇被不同地称为模式、基序、特征或指纹。这些基序的产生是因为蛋白质的特定区域(S)在结构和序列上都是保守的,这些区域对于蛋白质的结合特性或酶活性可能是重要的。这些结构要求对蛋白质序列中这些很小但很重要的部分(S)的进化施加了非常严格的限制。利用蛋白质序列模式来确定蛋白质的功能正迅速成为序列分析的基本工具之一。这一现实已经得到了许多作者的认可[1,2]。虽然已经对已发表的模式[3,4,5]进行了许多审查,但直到最近[6,7]才尝试系统地收集具有生物学意义的模式或发现新的模式。基于这些观察,我们在1988年决定积极开发一个模式数据库,用于搜索未知功能的序列。这个名为ProSite的数据库包含了一些在文献中发表的模式,但大多数模式是作者在过去四年中开发的。主导概念ProSite的设计遵循四个主要概念:完整性。为了使这样的汇编有助于确定蛋白质的功能,重要的是它包含尽可能多的具有生物学意义的模式。这些模式具有高度的特异性。在大多数情况下,我们选择了足够具体的模式,使得它们不会检测太多不相关的序列,但它们将检测大多数(如果不是全部)明显属于所考虑的集合的序列。
BACKGROUND PROSITE is a compilation of sites and patterns found in protein sequences; it can be used as a method ofdetermining thefunction of uncharacterized proteins translated from genomic or cDNA sequences. In some cases the sequence of an unknown protein is too distantly related to any protein of known structure to detect its resemblance by overall sequence alignment, but relationships can be revealed by the occurrence in its sequence of a particular cluster of residue types which is variously known as a pattern, motif, signature, or fingerprint. These motifs arise because specific region (s) of a protein which may be important, for example, for their binding properties or for their enzymatic activity are conserved in both structure and sequence. These structural requirements impose very tight constraints on the evolution of these small but important portion (s) of a protein sequence. The use of protein sequence patterns to determine the function of proteins is becoming very rapidly one of theessential tools of sequence analysis. This reality has been recognized by many authors [1, 2]. While there have been a number of reviews of published patterns [3, 4, 5], no attempt had been made until very recently [6, 7] to systematically collect biologically significant patterns or to discover new ones. Based on these observations, we decided in 1988, to actively pursue the development of a database of patterns which would be used to search against sequences of unknown function. This database, called PROSITE, contains some patterns which have been published in the literature, but the majority have been developed in the last four years by the author.LEADING CONCEPTS The design of PROSITE follows four leading concepts: Completeness. For such a compilation to be helpful in the determination of protein function, it is important that it contains as many biologically meaningful patterns as possible. High specificity of the patterns. In the majority of cases we have chosen patterns that are specific enough that they do not detect too many unrelated sequences, yet they will detect most, if not all, sequences that clearly belong to the set in consideration.