DATABASE OF HOMOLOGY-DERIVED PROTEIN STRUCTURES AND THE STRUCTURAL MEANING OF SEQUENCE ALIGNMENT

DATABASE OF HOMOLOGY-DERIVED PROTEIN STRUCTURES AND THE STRUCTURAL MEANING OF SEQUENCE ALIGNMENT
复制标题

DOI:
10.1002/prot.340090107
复制
发表时间:
1991-01-01
期刊:
PROTEINS-STRUCTURE FUNCTION AND GENETICS
影响因子:
--
通讯作者:
SCHNEIDER, R
SCHNEIDER, R
中科院分区:
其他
文献类型:
--
作者:
SANDER, C;SCHNEIDER, R

文献摘要

被引文献

相似文献

基于以下观察,通过使用序列同源性可以显着增加已知蛋白质三维结构的数据库。 (1) 已知序列数据库目前有超过 12,000 个蛋白质,比已知结构数据库大两个数量级。 (2)目前预测蛋白质结构最有力的方法是通过同源性建立模型。 (3)结构同源性可以从序列相似性水平来推断。 (4) 足以实现结构同源性的序列相似性阈值很大程度上取决于比对的长度。 在这里,我们首先通过对已知结构的蛋白质之间的比对进行详尽的调查来量化序列相似性、结构相似性和比对长度之间的关系,并报告作为比对长度的函数的同源阈值曲线。 然后,我们通过将基于阈值曲线视为同源的所有序列与已知结构的每个蛋白质进行比对,生成同源衍生的蛋白质二级结构 (HSSP) 数据库。 对于每个已知的蛋白质结构,衍生的数据库包含比对序列、二级结构、序列变异性和序列概况。 比对序列的三级结构是隐含的,但没有明确建模。 该数据库有效地将已知蛋白质结构的数量增加了五倍,达到 1800 以上。结果可用于评估序列数据库搜索中匹配的结构重要性、导出结构预测的偏好和模式、阐明保守残基的结构作用以及通过同源性建模三维细节。
The database of known protein three-dimensional structures can be significantly increased by the use of sequence homology, based on the following observations. (1) The database of known sequences, currently at more than 12,000 proteins, is two orders of magnitude larger than the database of known structures. (2) The currently most powerful method of predicting protein structures is model building by homology. (3) Structural homology can be inferred from the level of sequence similarity. (4) The threshold of sequence similarity sufficient for structural homology depends strongly on the length of the alignment. Here, we first quantify the relation between sequence similarity, structure similarity, and alignment length by an exhaustive survey of alignments between proteins of known structure and report a homology threshold curve as a function of alignment length. We then produce a database of homology-derived secondary structure of proteins (HSSP) by aligning to each protein of known structure all sequences deemed homologous on the basis of the threshold curve. For each known protein structure, the derived database contains the aligned sequences, secondary structure, sequence variability, and sequence profile. Tertiary structures of the aligned sequences are implied, but not modeled explicitly. The database effectively increases the number of known protein structures by a factor of five to more than 1800. The results may be useful in assessing the structural significance of matches in sequence database searches, in deriving preferences and patterns for structure prediction, in elucidating the structural role of conserved residues, and in modeling three-dimensional detail by homology.