CATH - a hierarchic classification of protein domain structures

CATH - a hierarchic classification of protein domain structures
复制标题

DOI:
10.1016/s0969-2126(97)00260-8
复制
发表时间:
1997-08-15
期刊:
影响因子:
5.7
通讯作者:
Thornton, JM
Thornton, JM
中科院分区:
生物学2区
文献类型:
--
作者:
Orengo, CA;Michie, AD;Thornton, JM

文献摘要

被引文献

相似文献

背景:蛋白质进化产生了结构相关的蛋白质家族,在这些家族中,序列同一性可能非常低。因此,基于结构的分类可以有效地识别已知结构中的意外关系,并且在最佳情况下也可以分配功能。已知蛋白质结构的数量不断增加,手动分类所有蛋白质的能力太大,因此,需要自动方法来快速评估蛋白质结构。结果:我们提出了一种半自动程序来推导新型蛋白质结构域分层分类(CATH)。我们的分类的四个主要水平是蛋白质类(C),架构(A),拓扑结构(T)和同源超家族(H)。类是最简单的级别,它本质上描述了每个域的二级结构组成。相比之下,建筑学总结了二级结构单元(如桶和三明治)的方向所揭示的形状。在拓扑级别,考虑顺序连接,使得相同架构的成员可能具有完全不同的拓扑。当结构属于相同的T-水平有适当的高相似性结合类似的功能,蛋白质被认为是进化相关的,并放入相同的同源superfamily.Conclusions:CATH产生的结构家族的分析揭示了蛋白质结构空间的突出特点。我们发现,近三分之一的同源超家族(H级)属于10个主要的T级,我们称之为超折叠,此外,近三分之二的这些H级集群成9个简单的架构。一个具有良好特征的蛋白质结构家族的数据库,如CATH,将有助于将结构-功能/进化关系分配给已知和新确定的蛋白质结构。
Background: Protein evolution gives rise to families of structurally related proteins, within which sequence identities can be extremely low. As a result, structure-based classifications can be effective at identifying unanticipated relationships in known structures and in optimal cases function can also be assigned. The ever increasing number of known protein structures is too large to classify all proteins manually, therefore, automatic methods are needed for fast evaluation of protein structures.Results: We present a semi-automatic procedure for deriving a novel hierarchical classification of protein domain structures (CATH). The four main levels of our classification are protein class (C), architecture (A), topology (T) and homologous superfamily (H). Class is the simplest level, and it essentially describes the secondary structure composition of each domain. In contrast, architecture summarises the shape revealed by the orientations of the secondary structure units, such as barrels and sandwiches. At the topology level, sequential connectivity is considered, such that members of the same architecture might have quite different topologies. When structures belonging to the same T-level have suitably high similarities combined with similar functions, the proteins are assumed to be evolutionarily related and put into the same homologous superfamily.Conclusions: Analysis of the structural families generated by CATH reveals the prominent features of protein structure space. We find that nearly a third of the homologous superfamilies (H-levels) belong to ten major T-levels, which we call superfolds, and furthermore that nearly two-thirds of these H-levels cluster into nine simple architectures. A database of well-characterised protein structure families, such as CATH, will facilitate the assignment of structure-function/evolution relationships to both known and newly determined protein structures.