THE PIR PROTEIN-SEQUENCE DATABASE

THE PIR PROTEIN-SEQUENCE DATABASE
复制标题

DOI:
10.1093/nar/19.suppl.2231
复制
发表时间:
1991-04-25
影响因子:
14.9
通讯作者:
GARAVELLI, JS
GARAVELLI, JS
中科院分区:
生物学2区
文献类型:
--
作者:
BARKER, WC;GEORGE, DG;GARAVELLI, JS

文献摘要

被引文献

相似文献

从1989年12月到1990年12月,蛋白质序列数据库从14,372条增加到26,798条,增加了86%;从3,977,903条增加到7,620,688条,增加了92%。这一增长很大程度上是由于引入了“第三节”。未验证条目,包含6,444个条目和1,816,684个残数。第二节。初步条目增加了4,785个条目,总数为12,607个条目和3,417,043个残数。第一节。注释和分类条目包含7,747个条目和2,386,941个残基。共有30,275次引用,18,257个来源。直接提交的引用占1653次。第1部分:注释和分类条目PIR的蛋白质序列数据库传统上一直被维护为参考数据汇编。这样的汇编包括每一项数据的一个“值”(或者可能是一个值范围),而不是再现个别研究报告的结果。我们的政策是通过将同一分子序列的多次测定数据组合到一个单独的条目中,以最大限度地减少冗余信息。形成数据库第1节的经过高度验证的非冗余数据收集支持需要代表性收集的统计分析,搜索效率高,并且为用户节省了验证、比较和组合来自各个报告的相关数据所需的大量时间。蛋白质序列数据库中完全注释条目的文本部分为评估数据库搜索结果提供了有用的信息。除了文献引用和关于序列实验测定的信息外,文本可能包括替代命名法,结构域结构,序列中的功能位点,遗传信息以及关于分子来源,功能,三级和四级结构以及物理化学性质的一般信息。该信息可用于在数据库中定位特定类型的蛋白质,或在发现两个蛋白质具有相似序列时比较其他特征。作为研究蛋白质进化的一组数据库的起源反映在数据的组织上,它基于蛋白质超家族的概念:一组蛋白质,其氨基酸序列可以显示为进化相关[4]。包含在特定超家族中的蛋白质意味着该蛋白质与该超家族中的其他蛋白质是同源的(从共同的祖先序列进化而来)。在Protein sequence Database的Section 1中,每个序列被分配为一组5个数字,其中第一个数字代表超家族。家族、亚家族和词条号将一个超家族细分为不同的蛋白质组,它们的差异分别超过50%、20%和5%。超族组织在解释数据库搜索时很有用,因为一个与数据库中已有序列真正同源的新序列应该显示出相应序列的相似性
From December 1989 to December 1990, the Protein Sequence Database grew from 14,372 to 26,798 entries, an 86% increase, and from 3,977,903 to 7,620,688 residues, a 92% increase. Much of this increase was due to the introduction of'Section 3. Unverified entries,'which contains 6,444 entries and 1,816,684 residues.'Section 2. Preliminary entries' increased by 4,785 entries to a total of 12,607 entries and 3,417,043 residues.'Section 1. Annotated and Classified Entries' contains 7,747 entries and 2,386,941 residues. There are altogether 30,275 citations to 18,257 sources. Direct submissions account for 1,653 of the citations.Section 1: Annotated and Classified Entries The Protein Sequence Database of the PIR has been traditionally maintained as a reference data compendium. Such a compendium includes one'value'for each item of data (or perhaps a range of values) rather than reproducing the results of individual research reports. Our policy has been to minimize redundant information by combining, into a single entry, data from multiple determinations of the sequence of the same molecule. The resulting highly verified, nonredundant data collection that forms Section 1 of the database supports statistical analyses that require a representative collection, is efficient to search, and saves users the considerable time needed to verify, compare, and combine related data from individual reports. The text portion of fully annotated entries of the Protein Sequence Database provides information useful for evaluating the results of database searches. In addition to literature citations and information about the experimental determination of the sequence, the text may include alternative nomenclature, domain structure, functional sites in the sequence, genetic information, and general information about the source, function, tertiary and quaternary structures, and physicochemical properties of the molecule. This information may be used to locate particular types of proteins in the database or to compare other features when two proteins are found to have similar sequences. The origin of the database as a set for the study of protein evolution is reflected in the organization of the data, which is based on the concept of a protein superfamily: a group of proteins whose amino acid sequences can be shown to be evolutionarily related [4]. Inclusion of a protein in a specific superfamily implies that the protein is homologous (evolved from a common ancestral sequence) with the other proteins in that superfamily. Each sequence in Section 1 of the Protein Sequence Database is assigned a set of five numbers, the first of which represents the superfamily. The family, subfamily and entry numbers subdivide a superfamily into groups of proteins that are more that 50%, 20% and 5% different, respectively. The superfamily ofganization is useful in interpreting database searches because a new sequence that istruly homologous with one already in the database should show sequence similarity to the corresponding