The SUPERFAMILY database in 2004: additions and improvements

The SUPERFAMILY database in 2004: additions and improvements
复制标题

DOI:
10.1093/nar/gkh117
复制
发表时间:
2004-01-01
影响因子:
14.9
通讯作者:
Gough, J
Gough, J
中科院分区:
生物学2区
文献类型:
--
作者:
Madera, M;Vogel, C;Gough, J

文献摘要

被引文献

相似文献

超家族数据库为蛋白质序列提供结构分配,并为结果分析。数据库的核心是一个配置文件库隐藏的马尔可夫模型,该模型代表所有已知结构的蛋白质。该库基于蛋白质的SCOP分类:每个模型对应于SCOP域,旨在表示整个超家族。我们已将库应用于所有完全测序的基因组(当前154),瑞士 - 普罗特和Trembl数据库以及其他序列集合的预测蛋白质。所有蛋白质的近60%至少有一项匹配,所有残留物的一半被任务覆盖。所有模型和完整结果均可在http://supfam.org上下载和在线浏览。用户可以研究其在所有完全测序的基因组中兴趣的超家族的分布,研究它结合的其他超家族并检索发生的蛋白质。另外,将集中在特定基因组上的整体上,首先可以找出其超家族组成,其次,将其与其他基因组进行比较以检测过度或代表性不足的超家族。此外,网络服务器提供以下标准服务:序列搜索;关键字搜索基因组,超家族和序列标识符;以及基因组,PDB和自定义序列的多重比对。
The SUPERFAMILY database provides structural assignments to protein sequences and a framework for analysis of the results. At the core of the database is a library of profile Hidden Markov Models that represent all proteins of known structure. The library is based on the SCOP classification of proteins: each model corresponds to a SCOP domain and aims to represent an entire superfamily. We have applied the library to predicted proteins from all completely sequenced genomes (currently 154), the Swiss-Prot and TrEMBL databases and other sequence collections. Close to 60% of all proteins have at least one match, and one half of all residues are covered by assignments. All models and full results are available for download and online browsing at http://supfam.org. Users can study the distribution of their superfamily of interest across all completely sequenced genomes, investigate with which other superfamilies it combines and retrieve proteins in which it occurs. Alternatively, concentrating on a particular genome as a whole, it is possible first, to find out its superfamily composition, and secondly, to compare it with that of other genomes to detect superfamilies that are over- or under-represented. In addition, the webserver provides the following standard services: sequence search; keyword search for genomes, superfamilies and sequence identifiers; and multiple alignment of genomic, PDB and custom sequences.