Touring protein fold space with Dali/FSSP

Touring protein fold space with Dali/FSSP
复制标题

DOI:
10.1093/nar/26.1.316
复制
发表时间:
1998-01-01
影响因子:
14.9
通讯作者:
Sander, C
Sander, C
中科院分区:
生物学2区
文献类型:
--
作者:
Holm, L;Sander, C

文献摘要

被引文献

相似文献

FSSP数据库及其新的补充——Dali结构域词典,呈现了对所有已知的3D蛋白质结构的持续更新的分类。该分类是通过一个自动结构比对程序(Dali)对蛋白质数据库中的结构进行两两比对得出的。从由此产生的结构近邻列举(其在折叠空间中形成令人惊讶的连续分布)中,我们通过三个步骤得出一个离散的折叠分类:(i)具有序列相关性的家族由一组有代表性的蛋白质链覆盖;(ii)基于结构模体的重复出现,蛋白质链被分解为结构域;(iii)折叠被定义为折叠空间中结构域的紧密簇。折叠分类、结构域定义以及用于序列 - 结构比对(穿线法)的测试集可在网站www.embl - ebi.ac.uk/dali上获取。该网络界面在折叠空间中的近邻之间、结构域和蛋白质之间以及结构和序列之间提供了丰富的链接网络,例如,可链接到一个处于序列相似性模糊区域的蛋白质家族明确多重比对数据库。蛋白质结构的Dali/FSSP组织提供了一幅当前已知的蛋白质宇宙区域的图谱,这对分析折叠原理、对蛋白质家族进行进化统一以及使从实验性结构测定中获得的信息回报最大化都很有用。
The FSSP database and its new supplement, the Dali Domain Dictionary, present a continuously updated classification of all known 3D protein structures, The classification is derived using an automatic structure alignment program (Dali) for the all-against-all comparison of structures in the Protein Data Bank. From the resulting enumeration of structural neighbours (which form a surprisingly continuous distribution in fold space) we derive a discrete fold classification in three steps: (i) sequence-related families are covered by a representative set of protein chains; (ii) protein chains are decomposed into structural domains based on the recurrence of structural motifs; (iii) folds are defined as tight clusters of domains in fold space, The fold classification, domain definitions and test sets for sequence-structure alignment (threading) are accessible on the web at www.embl-ebi.ac.uk/dali. The web interface provides a rich network of links between neighbours in fold space, between domains and proteins, and between structures and sequences leading, for example, to a database of explicit multiple alignments of protein families in the twilight zone of sequence similarity, The Dali/FSSP organization of protein structures provides a map of the currently known regions of the protein universe that is useful for the analysis of folding principles, for the evolutionary unification of protein families and for maximizing the information return from experimental structure determination.