Database support for 3D-protein data set analysis

Database support for 3D-protein data set analysis
复制标题

3D 蛋白质数据集分析的数据库支持

DOI:
--
复制
发表时间:
2003
期刊:
International Conference on Scientific and Statistical Database Management
影响因子:
--
通讯作者:
Wolfgang Lehner
Wolfgang Lehner
中科院分区:
--
文献类型:
--
作者:
Alexander Hinneburg;Wolfgang Lehner

文献摘要

参考文献

被引文献

相似文献

基因组研究的进展需要一个足够的基础设施来分析数据集。数据库系统反映了组织数据和加速分析过程的关键技术。本文以多维蛋白质数据库中频繁子结构的发现问题为基础,讨论了关系数据库系统的作用。具体的问题包括产生一组关联规则的频繁子结构与不同的长度和间隙之间的氨基酸残基的蛋白质。从数据库的角度来看,寻找关联规则的过程为更深入地分析数据材料奠定了基础,分为两个部分。第一部分通过计算给定的一组代表的最近邻来执行单个氨基酸残基的构象角空间的离散化。第二部分包括在适应一个著名的关联规则算法,以确定频繁的子结构。这个综合分析任务中的两个步骤都需要底层数据库的大量支持,以减少应用程序级别的编程开销。
The progress in genome research demands for an adequate infrastructure to analyze the data sets. Database systems reflect a key technology to organize data and speed up the analysis process. This paper discusses the role of a relational database system based on the problem of finding frequent substructures in multi-dimensional protein databases. The specific problem consists of producing a set of association rules regarding frequent substructures with different lengths and gaps between the amino acid residues of a protein. From a database point of view, the process of finding association rules building the base for a more in-depth analysis of the data material is split into two parts. The first part performs a discretization of the conformational angle space of a single amino acid residue by computing the nearest neighbor of a given set of representatives. The second part consists in adapting a well-known association rule algorithm to determine the frequent substructures. Both steps within this comprehensive analysis task requires substantial support of the underlying database in order to reduce the programming overhead at the application level.
DOI: 10.1006/jmbi.1997.0926
发表时间: 1997-04-18
影响因子: 5.6
作者:
Bower, MJ;Cohen, FE;Dunbrack, RL
通讯作者: Dunbrack, RL