Protein Families - Relating Protein Sequence, Structure, and Function

Protein Families - Relating Protein Sequence, Structure, and Function
复制标题

蛋白质家族 - 蛋白质序列、结构和功能的相关

DOI:
10.1002/9781118743089.ch4
复制
发表时间:
2013
期刊:
--
影响因子:
--
通讯作者:
Gough J
Gough J
中科院分区:
--
文献类型:
--
作者:
Gough J

文献摘要

相似文献

本章主要介绍SUPERFAMILY和Gene 3D,它们分别提供蛋白质序列结构域注释作为蛋白质结构分类(SCOP)和类、结构、拓扑、同源性(CATH)的配套资源。CATH和SCOP从蛋白质数据库(PDB)中获取3D原子分辨率结构,将它们划分为一个或多个域,并将它们分组为多级分类,例如家族和超家族。Gene 3D和SUPERFAMILY使用隐马尔可夫模型(HALGOT)对未知结构序列中的蛋白质结构域进行检测和分类。这两个资源都使用基于结构的结构域分类来创建已知和未知结构的同源序列的多重比对,代表每个超家族。这些网站包含数千万个序列的预先计算结果,并提供一个界面,用于单独或在基因组/系统发育背景下浏览、搜索、比较和可视化数据。
This chapter is primarily about SUPERFAMILY and Gene3D, which provide protein sequence domain annotations as companion resources to Structural Classification Of Proteins (SCOP), and Class, Architecture, Topology, Homology (CATH), respectively. CATH and SCOP take 3D atomic resolution structures from the Protein Data Bank (PDB), divide them into one or more domains and group them into a multi‐level classification including, for example, families and superfamilies. Gene3D and SUPERFAMILY detect and classify protein domains in sequences of unknown structure using hidden Markov models (HMMs). Both resources use the structure‐based domain classifications to create multiple alignments of homologous sequences of known and unknown structure, representing each superfamily. The websites contain precomputed results for tens of millions of sequences with an interface for browsing, searching, comparing and visualizing the data individually or in a genomic/phylogenetic context.