DWARF--a data warehouse system for analyzing protein families.

DWARF--a data warehouse system for analyzing protein families.
复制标题

DOI:
10.1186/1471-2105-7-495
复制
发表时间:
2006-11-09
期刊:
影响因子:
3
通讯作者:
Pleiss J
Pleiss J
中科院分区:
生物学4区
文献类型:
--
作者:
Fischer M;Thai QK;Grieb M;Pleiss J

文献摘要

参考文献

被引文献

相似文献

新兴的综合生物信息学领域提供了组织和系统分析大量高度多样化的生物数据的工具,从而可以对复杂的生物系统获得新的理解。数据仓库 DWARF 应用综合生物信息学方法来分析大型蛋白质家族。数据仓库系统 DWARF 集成了蛋白质折叠家族的序列、结构和功能注释数据。底层关系数据模型由三个主要部分组成,代表与蛋白质相关的实体(生化功能、来源生物、同源家族和超家族的分类)、蛋白质序列(位置特异性注释、突变信息)和蛋白质结构(二级结构信息、叠加的三级结构)。提供了从公共可用资源(ExPDB、GenBank、DSSP)中提取、转换和加载数据的工具来填充数据库。数据可以通过搜索和浏览界面以及对注释、序列或结构进行操作的分析工具来访问。我们将 DWARF 应用于 α/β-水解酶家族来托管脂肪酶工程数据库。 2.3版包含6138个序列和167个实验确定的蛋白质结构,它们被分配到37个超家族103个同源家族。 DWARF 旨在构建大型结构相关蛋白质家族的数据库,并通过序列、结构和功能注释的系统分析来评估其序列-结构-功能关系。它已被应用于根据序列预测生化特性,并成为蛋白质工程的重要工具。
The emerging field of integrative bioinformatics provides the tools to organize and systematically analyze vast amounts of highly diverse biological data and thus allows to gain a novel understanding of complex biological systems. The data warehouse DWARF applies integrative bioinformatics approaches to the analysis of large protein families. The data warehouse system DWARF integrates data on sequence, structure, and functional annotation for protein fold families. The underlying relational data model consists of three major sections representing entities related to the protein (biochemical function, source organism, classification to homologous families and superfamilies), the protein sequence (position-specific annotation, mutant information), and the protein structure (secondary structure information, superimposed tertiary structure). Tools for extracting, transforming and loading data from public available resources (ExPDB, GenBank, DSSP) are provided to populate the database. The data can be accessed by an interface for searching and browsing, and by analysis tools that operate on annotation, sequence, or structure. We applied DWARF to the family of α/β-hydrolases to host the Lipase Engineering database. Release 2.3 contains 6138 sequences and 167 experimentally determined protein structures, which are assigned to 37 superfamilies 103 homologous families. DWARF has been designed for constructing databases of large structurally related protein families and for evaluating their sequence-structure-function relationships by a systematic analysis of sequence, structure and functional annotation. It has been applied to predict biochemical properties from sequence, and serves as a valuable tool for protein engineering.
DOI: 10.1093/nar/gkg015
发表时间: 2003-01-01
影响因子: 14.9
作者:
Fischer, M;Pleiss, J
通讯作者: Pleiss, J
DOI: 10.1002/prot.20013
发表时间: 2004-06-01
影响因子: 2.9
作者:
Barth, S;Fischer, M;Pleiss, J
通讯作者: Pleiss, J
DOI: 10.1006/cbmr.1994.1011
发表时间: 1994-04-01
期刊: COMPUTERS AND BIOMEDICAL RESEARCH
影响因子: --
作者:
RITTER, O;KOCAB, P;SUHAI, S
通讯作者: SUHAI, S
DOI: 10.1016/s0923-2508(00)00121-2
发表时间: 2000-03-01
影响因子: 2.6
作者:
Schwede, T;Diemand, A;Peitsch, MC
通讯作者: Peitsch, MC
DOI: 10.1093/nar/gkh039
发表时间: 2004-01-01
影响因子: 14.9
作者:
Andreeva, A;Howorth, D;Murzin, AG
通讯作者: Murzin, AG