The CATH domain structure database.

The CATH domain structure database.
复制标题

CATH 域结构数据库。

DOI:
10.1002/0471721204.ch13
复制
发表时间:
2005
期刊:
Methods of biochemical analysis
影响因子:
--
通讯作者:
Janet M. Thornton
Janet M. Thornton
中科院分区:
--
文献类型:
--
作者:
C. Orengo;Frances M. G. Pearl;Janet M. Thornton

文献摘要

被引文献

相似文献

在进化过程中,蛋白质序列由于其残基的突变和残基的插入和缺失而发生变化。这些变化产生了相关蛋白质的家族,最早的蛋白质家族资源,仅仅基于序列数据,是在20世纪70年代由Dayhoff的开创性工作首次建立的。从那时起,建立了许多序列数据库,并且经常使用基于计算机科学领域的强大动态规划算法的比对方法来检测关系。这种方法非常有效地处理了发生在远亲进化中的残基插入和缺失。由于结构确定的技术挑战,结构数据一直比序列数据更稀疏。目前序列资源与构造资源之间存在两个数量级以上的差异。因此,当蛋白质数据库(PDB)包含大约16,000个结构条目时,国家生物技术信息中心(NCBI)的核苷酸序列数据库(Gen-Bank)包含超过1200万个条目。因此,虽然在20世纪70年代初就解决了第一个晶体结构,但直到20世纪90年代中期,结构分类才开始出现,主要是利用蛋白质结构分类(SCOP)(Murzin等,1995;Lo Conte, 2000), DALI (Holm和Sander, 1996)和CATH (Orengo等,1997;Pearl等,2001)数据库和数据资源(见表13.2)。此后出现了其他几种分类(例如,DDBASE (Sowdhamini et al., 1998), 3Dee (Dengler, Siddiqui, and Barton, 2001), DaliDD (Holm and Sander, 1998; Dietmann and Holm, 2001),详见Holm and Sander, 1994b和Orengo, 1994。这些数据库使用各种不同的算法来比较三维(3D)结构(参见第16章)。他们在测量相似度的方法上也有所不同
During evolution protein sequences change due to mutations in their residues and insertions and deletions of residues. These changes give rise to families of related proteins and the earliest protein family resources, based solely on sequence data, were first established in the 1970s by the pioneering work of Dayhoff. Since then many sequence databases have been established and relationships are often detected using alignment methods based on powerful dynamic programming algorithms adapted from the realm of computer science. Such methods handle the residue insertions and deletions occurring between distant evolutionary relatives very efficiently. The structural data has always been more sparse than the sequence data due to the technical challenges of structure determination. There is currently over two orders of magnitude discrepancy between the sequence and structure resources. Thus, while the Protein Data Bank (PDB) contains about 16,000 structural entries, the nucleotide sequence databank at the National Centre for Biotechnology Information (NCBI)(Gen-Bank) contains over 12 million entries.Therefore, although the first crystal structures were solved in the early 1970s, it was not until the mid-1990s that structural classifications began to emerge, primarily with Structural Classification of Proteins (SCOP)(Murzin et al., 1995; Lo Conte, 2000), DALI (Holm and Sander, 1996), and CATH (Orengo et al., 1997; Pearl et al., 2001) databases and data resources (see Table 13.2). Several other classifications have arisen since (see, for example, DDBASE (Sowdhamini et al., 1998), 3Dee (Dengler, Siddiqui, and Barton, 2001), DaliDD (Holm and Sander, 1998; Dietmann and Holm, 2001), reviewed in Holm and Sander, 1994b, and Orengo, 1994. These databases use a variety of different algorithms for comparing three-dimensional (3D) structures (see Chapter 16). They also differ in methods for measuring similarity between the