EFFICIENT DETECTION OF 3-DIMENSIONAL STRUCTURAL MOTIFS IN BIOLOGICAL MACROMOLECULES BY COMPUTER VISION TECHNIQUES

EFFICIENT DETECTION OF 3-DIMENSIONAL STRUCTURAL MOTIFS IN BIOLOGICAL MACROMOLECULES BY COMPUTER VISION TECHNIQUES
复制标题

DOI:
10.1073/pnas.88.23.10495
复制
发表时间:
1991-12-01
影响因子:
11.1
通讯作者:
WOLFSON, HJ
WOLFSON, HJ
中科院分区:
综合性期刊1区
文献类型:
--
作者:
NUSSINOV, R;WOLFSON, HJ

文献摘要

被引文献

相似文献

携带生物信息的大分子通常由包含重复结构基序的独立模块组成。检测蛋白质(或DNA)内的特定结构基序有助于阐明蛋白质(DNA元件)及其操作机制所起的作用。高分辨率的晶体学已知结构的数量正在迅速增加。然而,三维结构的比较是一个费力耗时的过程,通常需要手动阶段。到目前为止,还没有快速的自动化程序进行结构比较。我们提出了一个有效的O(n3)的最坏情况下的时间复杂度算法来实现这样的目标(其中n是在检查结构的原子数)。该方法是真正的三维,序列顺序无关,因此不敏感的差距,插入,或删除。该算法基于几何哈希范式,该范式最初是为计算机视觉中的对象识别问题而开发的。它介绍了一种基于变换不变表示的索引方法,特别是面向有效识别属于大型数据库的刚性对象中的部分结构。该算法适用于结构数据库的快速扫描,并将检测到一个先验未知的重复结构模体。该算法使用蛋白质(或DNA)结构,原子标签及其三维坐标。与结构有关的附加信息加速了比较。该算法是直接可并行化的,它的几个版本的计算机视觉应用程序已经实现了大规模并行连接机。一个原型版本的算法已经实现,并应用于蛋白质的子结构的检测。
Macromolecules carrying biological information often consist of independent modules containing recurring structural motifs. Detection of a specific structural motif within a protein (or DNA) aids in elucidating the role played by the protein (DNA element) and the mechanism of its operation. The number of crystallographically known structures at high resolution is increasing very rapidly. Yet, comparison of three-dimensional structures is a laborious time-consuming procedure that typically requires a manual phase. To date, there is no fast automated procedure for structural comparisons. We present an efficient O(n3) worst case time complexity algorithm for achieving such a goal (where n is the number of atoms in the examined structure). The method is truly three-dimensional, sequence-order-independent, and thus insensitive to gaps, insertions, or deletions. This algorithm is based on the geometric hashing paradigm, which was originally developed for object recognition problems in computer vision. It introduces an indexing approach based on transformation invariant representations and is especially geared toward efficient recognition of partial structures in rigid objects belonging to large data bases. This algorithm is suitable for quick scanning of structural data bases and will detect a recurring structural motif that is a priori unknown. The algorithm uses protein (or DNA) structures, atomic labels, and their three-dimensional coordinates. Additional information pertaining to the structure speeds the comparisons. The algorithm is straightforwardly parallelizable, and several versions of it for computer vision applications have been implemented on the massively parallel connection machine. A prototype version of the algorithm has been implemented and applied to the detection of substructures in proteins.