THE PIR PROTEIN-SEQUENCE DATABASE
THE PIR PROTEIN-SEQUENCE DATABASE
复制标题
DOI:
10.1093/nar/19.suppl.2231
复制
发表时间:
1991-04-25
影响因子:
14.9
通讯作者:
GARAVELLI, JS
中科院分区:
文献类型:
--
作者:
BARKER, WC;GEORGE, DG;GARAVELLI, JS
From December 1989 to December 1990, the Protein Sequence Database grew from 14,372 to 26,798 entries, an 86% increase, and from 3,977,903 to 7,620,688 residues, a 92% increase. Much of this increase was due to the introduction of'Section 3. Unverified entries,'which contains 6,444 entries and 1,816,684 residues.'Section 2. Preliminary entries' increased by 4,785 entries to a total of 12,607 entries and 3,417,043 residues.'Section 1. Annotated and Classified Entries' contains 7,747 entries and 2,386,941 residues. There are altogether 30,275 citations to 18,257 sources. Direct submissions account for 1,653 of the citations.Section 1: Annotated and Classified Entries The Protein Sequence Database of the PIR has been traditionally maintained as a reference data compendium. Such a compendium includes one'value'for each item of data (or perhaps a range of values) rather than reproducing the results of individual research reports. Our policy has been to minimize redundant information by combining, into a single entry, data from multiple determinations of the sequence of the same molecule. The resulting highly verified, nonredundant data collection that forms Section 1 of the database supports statistical analyses that require a representative collection, is efficient to search, and saves users the considerable time needed to verify, compare, and combine related data from individual reports. The text portion of fully annotated entries of the Protein Sequence Database provides information useful for evaluating the results of database searches. In addition to literature citations and information about the experimental determination of the sequence, the text may include alternative nomenclature, domain structure, functional sites in the sequence, genetic information, and general information about the source, function, tertiary and quaternary structures, and physicochemical properties of the molecule. This information may be used to locate particular types of proteins in the database or to compare other features when two proteins are found to have similar sequences. The origin of the database as a set for the study of protein evolution is reflected in the organization of the data, which is based on the concept of a protein superfamily: a group of proteins whose amino acid sequences can be shown to be evolutionarily related [4]. Inclusion of a protein in a specific superfamily implies that the protein is homologous (evolved from a common ancestral sequence) with the other proteins in that superfamily. Each sequence in Section 1 of the Protein Sequence Database is assigned a set of five numbers, the first of which represents the superfamily. The family, subfamily and entry numbers subdivide a superfamily into groups of proteins that are more that 50%, 20% and 5% different, respectively. The superfamily ofganization is useful in interpreting database searches because a new sequence that istruly homologous with one already in the database should show sequence similarity to the corresponding